---
title: "Throughput and latency"
description: "One queue of 200 partitions offered 300,000 to 4,000,000 msg/s on three nodes, with Kafka, Redpanda and Pulsar on the same machines: latency at each rate, each system's ceiling, CPU, and where Queen's latency comes from."
---

> Queen MQ documentation, for AI agents
> Complete self-contained summary of Queen MQ: https://queenmq.com/llms-brief.txt
> Fetch that first when the question is about the product rather than about this page.
> Index of all pages: https://queenmq.com/llms.txt

# Throughput and latency

Two hundred partitions in one queue is the shape that suits a broker with a log per partition
best: few logs, full batches, little state to keep. So that is where we raised the rate, from
300,000 to 4,000,000 messages a second, to find each system's ceiling and to see what latency it
paid on the way up. Queen holds a million messages a second here, and it is not the fastest system
on this page.

## The shape

One queue, or one topic, of 200 partitions. Producers pushed 256-byte JSON messages, 100 at a time
to one partition, and consumers took up to 1,000 messages per pop or fetch and acknowledged them
asynchronously. Each point ran for 60 seconds and the figures cover its last 30. Queen
2.0.0-beta.2 ran on 2026-10-01 with deduplication off and pops of up to 10 partitions; Kafka 4.3.1
ran on 2026-09-30, Redpanda 26.2.3 and Pulsar 4.2.4 on 2026-10-01, on the same brokers and load
machines. Runs and scripts: `benchmark-queen/2026-09-30-stage03-3node/` (Queen, tags `mb2c-q1-t200-*`)
and `benchmark-queen/2026-09-30-kafka-pulsar/runs/` (the others).

## Latency by rate

End-to-end latency, p50 / p99 in milliseconds, the worst of the load processes. "Behind" marks a
point where the system consumed under 97% of the offered rate or its p99 passed 5 seconds.

| Offered msg/s | Queen | Kafka | Redpanda | Pulsar |
|---|---|---|---|---|
| 300,000 | 11 / 22 | 8 / 10 | 9 / 66 | 7 / 15 |
| 600,000 | 19 / 98 | 8 / 13 | 16 / 181 | 7 / 18 |
| 1,000,000 | 61 / 146 | 8 / 18 | 16 / 187 | 7 / 21 |
| 1,500,000 | behind: 1.34M out, p50 4.6 s | 8 / 35 | 32 / 322 | 8 / 20 |
| 2,000,000 | behind: 1.50M in, 1.23M out | 9 / 59 | 48 / 253 | 12 / 21 |
| 3,000,000 | not run | 9 / 74 | behind: 2.92M out, p99 23 s | not run |
| 4,000,000 | not run | 17 / 75 | behind: p99 19 s | not run |

**Figure.** p99 end-to-end latency against the offered rate, one queue or topic of 200 partitions on three nodes, log scale. Kafka goes from 10 ms at 300,000 msg/s to 75 ms at 4,000,000. Pulsar stays between 15 and 21 ms up to 2,000,000. Queen goes from 22 ms at 300,000 to 98 ms at 600,000 and 146 ms at 1,000,000, and falls behind at 1,500,000. Redpanda goes from 66 ms to between 181 and 322 ms up to 2,000,000 and falls behind at 3,000,000.

Each line stops at the last rate the system kept up with. In this shape Kafka's latency is the one to beat, and Queen's ceiling is about 1.5M msg/s for one queue, set by one leader.

Values: p99 end-to-end latency.

| offered msg/s | Queen | Kafka | Redpanda | Pulsar |
|---|---|---|---|---|
| 300k | 22 ms | 10 ms | 66 ms | 15 ms |
| 600k | 98 ms | 13 ms | 181 ms | 18 ms |
| 1M | 146 ms | 18 ms | 187 ms | 21 ms |
| 1.5M |  | 35 ms | 322 ms | 20 ms |
| 2M |  | 59 ms | 253 ms | 21 ms |
| 3M |  | 74 ms |  |  |
| 4M |  | 75 ms |  |  |

Source: `benchmark-queen/2026-09-30-stage03-3node (Queen), benchmark-queen/2026-09-30-kafka-pulsar/runs (the others)`.

What each system had done when it answered a producer:

| System | Copies | Acknowledged after |
|---|---|---|
| Queen | 3 | the entry is fsynced on 2 of the 3 nodes |
| Kafka | 3 (`acks=all`, `min.insync.replicas=2`) | every in-sync replica has the batch, in page cache; no fsync |
| Redpanda | 3 | 2 of 3 have fsynced (`write_caching` off, its default) |
| Pulsar | 3 (ensemble 3, write quorum 3, ack quorum 2) | 2 bookies have fsynced their journal |

Kafka's line is the one to beat: 8 ms at the median from 300,000 to 3,000,000 msg/s. Pulsar waits
for the same two fsyncs as Queen and stays as flat up to 2M. Queen holds a p99 of 22 ms at 300,000
and climbs from there, to 98 ms at 600,000 and 146 ms at a million. Redpanda's median stays low and
its p99 sits near 200 ms. Its README in the archive traces that tail to one broker, the one with CPU
steal and the slowest disk, and notes that Redpanda is tested on XFS while this rig has one ext4
disk per broker.

## The ceiling

At 1,500,000 msg/s offered, Queen accepted everything but its consumers took 1.34M a second, and
the backlog pushed the median e2e to 4.6 seconds. At 2M it accepted 1.5M and shed the rest at the
producers' in-flight cap. The leader used 12 to 13 of its 16 cores at both points. That is
Queen's ceiling for one queue: about 1.5M msg/s pushed, set by one leader. Kafka carried 4M msg/s
in and out at a p99 of 75 ms, and we did not push it further. Pulsar carried 2M at 21 ms and was not
run higher. Redpanda fell behind at 3M.

**Figure.** Messages consumed per second against messages offered, one queue or topic of 200 partitions on three nodes. Kafka consumes everything up to 4,000,000 msg/s. Pulsar consumes everything up to 2,000,000, the highest rate it ran. Redpanda keeps up to 2,000,000 and consumes 2.92M of 3M. Queen keeps up to 1,000,000, consumes 1.34M of 1.5M and 1.23M of 2M.

Where each line leaves the diagonal, the system stopped keeping up. Queen's ceiling for one queue is about 1.5M msg/s pushed, with its leader at 12 to 13 of 16 cores.

Values: consumed msg/s.

| offered msg/s | Queen | Kafka | Redpanda | Pulsar |
|---|---|---|---|---|
| 300k | 300k | 300k | 300k | 300k |
| 600k | 600k | 600k | 600k | 600k |
| 1M | 1M | 1M | 1M | 1M |
| 1.5M | 1.34M (see caption) | 1.5M | 1.5M | 1.5M |
| 2M | 1.23M (see caption) | 2M | 2M | 2M |
| 3M |  | 3M | 2.92M (see caption) |  |
| 4M |  | 4M |  |  |

Source: `benchmark-queen/2026-09-30-stage03-3node (Queen), benchmark-queen/2026-09-30-kafka-pulsar/runs (the others)`.

## Where Queen's latency comes from

A push is answered once its entry is fsynced on two of the three nodes, as in Pulsar. Two more
things are Queen's own, and both are choices we would make again.

A pop is a write. The consumption engine on the leader hands out a lease, and the answer waits
until the checkpoint carrying that lease has committed through the log, which happens every 5 ms
(`QUEEN_CONSUME_CHECKPOINT_MS`). A new leader therefore knows every lease the old one gave, and no
consumer is ever handed a message that another consumer still holds. A message pays for two
replicated commits between its push and its delivery. `QUEEN_CONSUME_FAST=1` answers before the
checkpoint and gives that guarantee up; no run here used it.

One leader orders every write of a raft group. That leader is what makes a transaction one entry in
one log, and it is also where the work queues up as the rate climbs: at 1M msg/s it used 9 of its 16
cores while the two followers used 3.7 and 4.7. Kafka, Redpanda and Pulsar spread their partition
leaders over all three brokers. A cluster can run several raft groups (`QUEEN_RAFT_GROUPS`), each
with its own leader, and a tenant lives in exactly one. Every run on this page used one tenant, so
one group and one leader.

## CPU

Cores used by each system's own processes, summed over the three brokers (48 vCPU in all).

| Offered msg/s | Queen | Kafka | Redpanda | Pulsar |
|---|---|---|---|---|
| 300,000 | 9.0 | 3.6 | 4.1 | 2.7 |
| 600,000 | 13.0 | 6.3 | 8.4 | 4.5 |
| 1,000,000 | 17.2 | 9.7 | 14.2 | 7.1 |
| 1,500,000 | 24.9 (fell behind) | 12.6 | 19.7 | 10.0 |
| 2,000,000 | 27.2 (fell behind) | 14.5 | 23.6 | 12.7 |
| 3,000,000 | not run | 18.4 | 31.9 (fell behind) | not run |
| 4,000,000 | not run | 19.4 | 37.7 (fell behind) | not run |

**Figure.** Cores used by each system against the offered rate, summed over three brokers with 48 vCPU in all, one queue or topic of 200 partitions. At 1,000,000 msg/s Queen used 17.2 cores, Redpanda 14.2, Kafka 9.7 and Pulsar 7.1. Queen reached 24.9 and 27.2 cores at 1.5M and 2M, where it had fallen behind; Kafka used 19.4 cores at 4M.

In this shape, few partitions and a high rate, Queen spends two to three times the CPU of Kafka and Pulsar per message. Hollow marks are rates the system no longer kept up with.

Values: cores, three brokers.

| offered msg/s | Queen | Kafka | Redpanda | Pulsar |
|---|---|---|---|---|
| 300k | 9 | 4 | 4 | 3 |
| 600k | 13 | 6 | 8 | 5 |
| 1M | 17 | 10 | 14 | 7 |
| 1.5M | 25 (see caption) | 13 | 20 | 10 |
| 2M | 27 (see caption) | 15 | 24 | 13 |
| 3M |  | 18 | 32 (see caption) |  |
| 4M |  | 19 | 38 (see caption) |  |

Source: `benchmark-queen/2026-09-30-stage03-3node (Queen), benchmark-queen/2026-09-30-kafka-pulsar/runs (the others)`.

Redpanda's column is its own count of busy shard time. Its reactors poll, so the operating system
counts more: 39 cores at 1M msg/s. In this shape Queen spends two to three times the CPU of Kafka
and Pulsar for each message. At 100,000 partitions the picture turns around; see
[partition count](/benchmarks/partitions/).

## Deduplication

These runs had deduplication off, because Kafka, Redpanda and Pulsar have no equivalent to price
in. An earlier build measured what it costs: on 2026-09-29, raft at `e3456a48`, one queue of
500,000 partitions at 1M msg/s ran at an e2e p99 of 249 to 272 ms with a 60-second dedup window and
232 to 257 ms without one, in 90-second runs (`benchmark-queen/2026-09-29-1m-3node/`). That build
predates the consumption engine, so take the difference from it and the latency from the table
above. The [soak](/benchmarks/soak/) runs with deduplication on.

## When every message has its own partition

Every run on this page pushes 100 messages to one partition per request, and the broker's work
grows with the partitions a command touches more than with the messages it carries. When each
push carries one message per partition, the same broker does far more work per message. The last
measurement of that shape had each push carry 10 messages for 10 different partitions, over 10
queues of 100,000 partitions, and peaked near 114,000 msg/s pushed (2026-09-30, the Stage 0-3
build committed as `08050fed`, run `s08-mk10-q10-t1m-r200k`). That build predates the consumption
engine and the shape has not been measured since. Plan for it if your producers send one message
per entity at a time.

## What this page does not show

A million a second for long. Each point ran 60 seconds, and the longest run at 1M msg/s on the
current code is ten minutes of Kafka clients through the facade, which had stalls (see
[Kafka clients](/benchmarks/kafka-clients/)); our runs of hours are at 500,000 msg/s. Nor does it
show more than one raft group, payloads other than 256 bytes, or consumers that do any work. The
figures are 2.0.0-beta.2: the push path changed in beta.5, and Queen's own clients have not been
measured with four load machines since.

Source: https://queenmq.com/benchmarks/throughput/index.mdx
