A queue is a named set of partitions, and a partition is the ordered stream of one entity: a customer, an order, a device. You never declare either of them. The first push that names a partition creates it, in the same log entry that stores the message, so you can key your data by what it is instead of by a shard count you had to guess before the first customer signed up.
const res = await client
.queue('orders')
.partition('customer-42')
.push([{ data: { orderId: 9137, amount: 99.5 } }])curl -s -X POST http://localhost:6632/api/v1/push -H 'content-type: application/json' -d '{
"items": [{ "queue": "orders", "partition": "customer-42",
"transactionId": "order-9137-created", "payload": { "orderId": 9137 } }]
}'[{"index":0,"message_id":"01a0fcd5-2cdf-7000-a9bb-2f832df96c36","transaction_id":"order-9137-created","queueName":"orders","status":"queued","offset":0}]The answer is HTTP 201 with one verdict per item, in the order you sent them. status is
queued (stored at that offset), duplicate (the partition already has this transactionId,
nothing was written, see dedup) or error (refused, nothing written, and no
offset). The 201 says the request was taken, so read the status of each item.
Why one partition per entity
Most brokers order messages per shard. You pick a number of partitions when you create a topic, your keys are hashed onto them, and every entity shares its shard with whoever else hashed there. A slow customer holds up the customers behind it in the same shard, and changing the count later moves keys from one shard to another. So you guess the count early, and you live with the guess.
We made the partition the entity because in Queen it costs almost nothing. A partition is a few rows in the broker’s state, about 2 KB of the leader’s memory, plus one cursor for each consumer group that reads it; it is not a file, a process or a replica set. That is why there is no count to plan and nothing to provision, and why a queue can hold millions of them: on three 16-vCPU nodes, one queue of ten million partitions carried 1,000,000 messages a second, pushed and consumed, at an end-to-end p99 of 122 ms (2.0.0-beta.2, 2026-10-01; see partition count).
Choose the partition key at the boundary where order matters to you. If a customer’s events must be handled in order, the partition is the customer. A partition per message would give you no order to keep and nothing to gain.
Order and parallelism
Offsets in a partition start at 0 and have no gaps, and every reader sees the same order. A
consumer group holds one leased batch per partition at a time, which is what keeps that order
while the group’s workers run in parallel: two workers never hold the same partition at once, so
they work on different entities. A slow or stuck partition delays only itself, as far as the
broker is concerned. A client can still couple them: by default one pop may lease several
partitions, and a worker that handles them one after another makes the second wait for the first.
.partitions(1) keeps each pop to one partition, and the webhooks example
shows the difference with a dead endpoint listed first.
The other side of that rule is that one hot partition is sequential, by design. Twenty workers on one busy partition do the work of one, and twenty partitions with work keep twenty workers busy. Parallelism comes from the number of entities with something to do.
A push that names no partition goes to the partition Default. One push request may name several
queues and partitions, and each item gets its own verdict, so a plain push is not atomic across
partitions. When several messages must land together, or not at all, push them in one
transaction.
Queue options
A queue created by a push starts with the defaults below. POST /api/v1/configure creates a queue
or changes one, and merges by default: an option you leave out keeps its stored value
("mode": "replace" resets the omitted ones instead).
curl -s -X POST http://localhost:6632/api/v1/configure -H 'content-type: application/json' \
-d '{"queue":"orders","options":{"leaseTime":30,"retryLimit":5,"dedupWindowSeconds":86400}}'| Option | Default | Effect |
|---|---|---|
leaseTime |
60 s (300 s if /configure created the queue) |
How long a pop’s lease lasts unless the pop sends leaseSeconds. |
retryLimit |
3 | How many failed acks redeliver a message. The next failed files it in the dead-letter queue. |
deadLetterQueue, dlqAfterMaxRetries |
true |
File a message that ran out of retries in the DLQ. With both off, drop it and move on. |
dedupWindowSeconds |
3600 | How long a partition remembers transactionIds. 0 turns dedup off. |
delayedProcessing |
0 | Seconds after its push before a message can be delivered. |
windowBuffer |
0 | A partition delivers nothing until it has been quiet this many seconds, so a burst arrives as one batch. |
retentionEnabled |
false |
Must be true for the two options below to act. |
retentionSeconds |
0 | Remove messages older than this, consumed or not. |
completedRetentionSeconds |
0 | Remove messages older than this that every group’s cursor has passed. A partition no group reads keeps them. |
maxWaitTimeSeconds |
0 | Remove messages older than this, consumed or not, with no retentionEnabled needed. |
encryptionEnabled |
false |
Encrypt payloads at rest. Needs QUEEN_ENCRYPTION_KEY (64 hex characters) on every node. |
namespace, task |
from the name | Labels a discovery pop matches on. |
A queue created by a push takes its labels from the first two dotted parts of its name:
billing.invoices gets namespace billing and task invoices, and an undotted orders gets
namespace orders and no task. A queue created by /configure has no labels unless you send them.
priority, retryDelay, ttl, maxSize and minPopWaitTime are accepted and stored, and have
no effect in 2.0.
Replay from retained history
An ack moves a consumer group’s cursor, and that is all it does: the message stays in its partition until retention removes it. Any group can read it again, so a new projection, an audit or a fixed bug can replay what already happened without asking anyone to send it twice. A new group that starts from the beginning reads the retained history without touching any other group’s cursor:
// A fresh group with subscriptionMode('all') reads the whole retained
// lane from the beginning, without touching any other group's cursor.
const replayed = await client
.queue('orders')
.partition('customer-91')
.group('audit')
.subscriptionMode('all')
.batch(100)
.wait(true)
.pop()subscriptionFrom('2026-09-01T00:00:00Z') starts a new group at a point in time instead, and an
existing group moves with a seek. Both are on consuming.
Ephemeral queues
Some traffic is not worth a disk write: typing indicators, presence, progress bars. Ephemeral
queues (/api/v1/ephemeral/*) keep their messages in memory, with the same partitions, consumer
groups, leases and acks. In a cluster each ephemeral partition has one owner node, which the
nodes agree on by hashing over the live members, and a request that lands elsewhere is forwarded
to it. Their contents are not in the replicated log, so a crashed node loses what it held. See
ephemeral queues.
Limits
- A partition is sequential for each group, and two partitions have no order between them.
- Every node holds the state of every partition, so memory sets how many a cluster can hold: ten million took two thirds of a 31 GB node in the run above.
- Retention is off by default. Every message is kept, and the disk grows, until you set
retentionEnabledwith a retention option, ormaxWaitTimeSeconds. At 85% disk a node refuses writes that grow storage (507). - An encrypted queue on a node without
QUEEN_ENCRYPTION_KEYstores plaintext, silently: the push is not refused and nothing is logged. Dead letters of an encrypted queue are stored decrypted. - A partition that retention has emptied is deleted after 30 days without a write, a lease or a
dead letter (
PARTITION_CLEANUP_DAYS). - Names share a budget of 447 bytes: the tenant, the queue and the longer of the partition or consumer group name. The tenant counts even with tenancy off (the default tenant’s name is 36 bytes), so a queue name and a partition name have 411 bytes between them. Past it, the call answers 413.
- One request body is at most 64 MiB (
QUEEN_MAX_BODY_BYTES), and one planned entry at most 96 MiB (QUEEN_RAFT_ENTRY_MAX_BYTES). A push group past the entry limit is refused item by item, witherrorinside the 201.