Before 2.0.0 (2026-10-03), Queen MQ kept its data in PostgreSQL. 2.0 replaced that with its own log, replicated with Raft, on each node’s local disk, and a 2.0 node does not read a 1.x database. This page lists every release since 2.0.0, newest first. The 1.x entries are in CHANGELOG.md on GitHub.
2.1.0
Server: a standby cluster. A second cluster can now replay the first one’s log and take over
when the first is lost. The standby’s leader reads the source’s committed entries over the raft
port and proposes each one into its own log, so the standby holds the source’s messages, cursors,
leases, KV, timers and dedup window, a moment behind. It answers reads, refuses
writes with 503 and "code": "standby" on every node, and becomes an ordinary cluster with
POST /api/v1/system/link/promote. An empty standby follows a young source from the start of its
log (QUEEN_LINK_STANDBY); a source with a history seeds the standby’s first node with its
snapshot (QUEEN_LINK_SEED). Replication is asynchronous: a promotion after a crash loses at most
what the standby had not read yet, and a planned switch loses nothing. On one laptop, with both
clusters and the load generator sharing a disk, a standby stayed within 0.25 s of a source taking
200,000 messages a second. The source keeps its log for the standby on every node and across
restarts, and gives it up when its own disk fills (QUEEN_LINK_HOLD_S,
QUEEN_LINK_HOLD_DISK_PCT). GET /api/v1/system/link, the link block of /health and the
queen_link_* series say where a standby is. The source needs QUEEN_LINK_TOKEN, the standby
QUEEN_LINK_SOURCE and QUEEN_LINK_SOURCE_TOKEN; a cluster with none of them behaves as before.
See Standby cluster.
Server: the last node of a cluster to stop no longer hangs. A node that led a cluster whose
other nodes were already down, with an entry in its log that could no longer commit, waited for
that entry for ever while it shut down, and used a whole core doing it. A stopping node now leaves
once every caller of its entries in flight has an answer; an entry that timed out had already
answered retry.
Server: a leader that no majority has acknowledged for two seconds no longer stops for good. A leader cut off for two or three seconds and not replaced, or one whose majority needs a follower with a slow disk, had its next write refused by raft, took the refusal for a lost leadership and stopped planning. It still led, so nothing started it again: every request ran into its deadline until the leadership changed or the node restarted. Every release since 2.0.0 has it. Such a leader now keeps its entries and logs them, in order, when a majority answers again, and a leader that finds itself stopped while it leads starts again after 5 s. Found by Jepsen on a five-node cluster with two slow disks. With the leader of five nodes cut off for 2.4 s, the build before the fix took no write afterwards in 11 of the 12 trials where the leader kept its leadership, and this one took the next write in all 10. The log lines that say when it happens are on Monitoring.
Server: locks and semaphores. A lock is a lease: one holder at a time, for a lifetime the holder
declares and renews, with a token that fences a holder that outlived it. A semaphore is the same
lease with up to 1,024 permits. One route, POST /api/v1/locks, carries acquire, renew,
release and get, up to 64 locks a call, and answers a verdict per operation with HTTP 200
(acquired: false, reason: "held" is not an error). A permit is one KV row in the namespace
queen-locks (key <name>#<slot>, the owner as its value, the lifetime as its TTL), so the route
adds nothing to the log’s format: an acquire is a putIfAbsent, a renew a put with expect, a
release a delete with expect. The token is the row’s version and changes at every renew.
An owner makes a call safe to send again: the same owner is answered the permit it has. Locks
count against the tenant’s KV quota and KV write rate, need a read-write token, and behind the
proxy are part of the KV plan and never quota-blocked. New metrics: queen_locks_ops_total and
queen_locks_op_duration_milliseconds. A lock expires; a holder that crashed keeps it until its
lifetime ends, a waiter polls, and there are no read/write locks.
Server: KV check, a precondition that writes nothing. {"op":"check","ns","key","expect"} asserts
that a key is at a version, or absent with expect: 0. With required: true it gates a KV call or
a transaction on a key the call does not write. It is what guards a transaction with a lock: every
acquire and renew answers a guard, a check of the permit’s row at its token, and a transaction
that carries it commits only while the lock is still the caller’s.
Server: KV versions are a fencing token. On one key, a later write always has a higher version than every earlier one, across a delete and a re-create, an expiry, a restart and a change of leader. The planner always assigned them that way; it is now the documented contract, with a test, and the docs no longer say not to rely on their order.
Clients: lock, semaphore, check and guard, in all six SDKs. queen.lock(name) and
queen.semaphore(name, limit), each with a lifetime, return a handle that acquires (with an
optional wait), keeps the current token, renews and releases; transaction().guard(lock) puts the guard in the
commit and reads the token when the commit is sent. The JavaScript, Python, Go and Rust handles
renew in the background every third of the lifetime and signal a lost lock. The PHP and C++
handles have no background renewal: keepAlive() (keep_alive() in C++) renews at a checkpoint
in the work loop. A guarded commit that lost only to its own handle’s renewal is sent again with
the new token. kv.check is on the KV client and the transaction’s KV builder of each SDK. In the
JavaScript, Python, Go, Rust and C++ clients 2.1.0 and the PHP client 2.4.0.
Dashboard: a Locks page. Every held lock and semaphore permit with its holder, since when it
is held, its last renewal and when it expires. It reads the permits as KV rows, so a viewer can
open it, and it writes nothing: the drawer shows the guard a transaction carries and the call
that releases the lease period on screen. get answers the same since for each holder: a
renewal does not move it.
Server: a KV call that mixed a read with writes that all lost no longer hangs. A batch on
POST /api/v1/kv such as a putIfAbsent that lost beside a get wrote nothing, so it had no log
position, and its read waited for one until the statement timeout (30 s) and answered 503. Such a
call now takes a position of its own and answers at once.
Server: the upgrade is one-way. A node forwards a check to the leader only when every member can
read it. Once every member runs this release the leader raises the cluster version to 5, and
from then on 2.0.4 or older refuses to start on the data directory. Until then a call that carries a check answers 503
(kv_check_needs_cluster_version_5); acquire, renew and release need no new format and work as
soon as the node you call runs this release. To keep the way back open while the release bakes,
set QUEEN_RAFT_CLUSTER_VERSION_MS=0.
2.0.4 - 2026-10-08
Dashboard, sign-in page and docs: the logo is a yellow sunflower. A new drawing replaces the red one of 2.0.3: seven yellow petals round the q of “queen”, with “message queue” under the name. The dashboard shows the q and its petals beside the word “queen”, the sign-in page and the README show the whole logo, and the tab icon is the q with its petals. The picture of the bee and the sunflower on the documentation site takes the same yellow.
Dashboard: the logo’s yellow is the accent. Primary buttons are yellow with black labels, and the page you are on, selected segments and focus rings take the same colour. On light surfaces, text and rings in that colour are a darker gold, so they can be read. The sign-in page and the documentation site follow. Status colours keep their meanings: amber is attention, coral is failure.
Dashboard: System, Light and Dark. The theme toggle in the top bar is now a choice of three. System follows the device’s colour scheme as it changes, and is what a browser with no stored choice gets. Light and Dark stay fixed on that device until System is chosen again, and a choice made before 2.0.4 is kept.
Dashboard: supervisor cards share one grid. The Supervisors page gave each publication group its own two-column grid, so two groups with one instance each took two rows and left the second column empty. The instances now share one grid, each card under its group’s name: two to a row on a wide screen, one on a narrow one.
The broker’s engine, its API and the clients are as in 2.0.3.
2.0.3 - 2026-10-08
Server: a pop that nobody receives no longer counts a delivery attempt. A pop’s answer waits for
the checkpoint that holds its lease. When that takes longer than the margin the broker keeps before
the request’s deadline, or the client leaves, nobody receives the answer, and the broker hands the
leases back at once. It left the delivery attempt counted all the same, so the next consumer got
the message as a redelivery. With deliveryAttempt 2 on its first real delivery, a Laravel job
with tries = 1 failed with “attempted too many times” without ever running. In a 23-hour run of
4.1 million jobs at 50 jobs/s, 4 jobs failed that way, each about 0.4 s after its push, while every
message went out to a worker once; checkpoint writes reached 0.45 s against a 250 ms margin. Handing
back an unreceived claim now also takes back its attempt, on the leader’s engine and through the
Nack a follower sends. A lease that really expired still counts.
Clients: a consumer can report to the dashboard, off by default. Every SDK’s consumer builder
takes a supervision option that names a group for the dashboard: .supervision({ group }) in
JavaScript, Supervision(&queen.SupervisionConfig{...}) in Go, supervision(...) in the Python,
Rust, C++ and PHP clients. Each consume call then writes an observation to the broker’s
queen-supervisor KV namespace every 10 seconds, and a last one when it stops: its execution model,
its live loops against the configured concurrency, the handlers running, the handler calls completed
and failed, and the age of the oldest handler still running. No payload and no error text is
published. Publishing is best effort and bounded, and it needs a credential that can write KV. It only
observes: it restarts nothing, scales nothing and changes no ack or lease, and with the option off
there is no timer and no request. The PHP client publishes at its cooperative checkpoints, so a long
synchronous handler leaves its last observation stale. In the JavaScript client 2.0.4, the Python, Go,
Rust and C++ clients 2.0.3 and the PHP client 2.3.1.
CLI: queenctl tail --supervision-group. tail reports the same way when the flag names a
group, one identity for each invocation.
Dashboard: the Supervisors page shows the consumers that report. Beside the process supervisors it lists each reporting consumer with its execution model (async tasks, goroutines, threads or cooperative loops), its live loops against the configured concurrency, its busy handlers, its completed and failed handler calls, its last completion and the age of its oldest handler in flight. A completed handler does not prove its ack succeeded, and a fresh heartbeat does not prove progress; a stale or inconsistent report stays unconfirmed. Process budgets and readiness stay with the process supervisors.
Dashboard, sign-in page and docs: a new logo. The q that is a sunflower replaces the sunflower and the bee: as a badge beside the word “queen” in the dashboard, as the wordmark on the sign-in page and in the README, and as the tab icon. The sign-in page now takes the dashboard’s light or dark scheme.
JS client 2.0.3 - 2026-10-07
JS client: stopping a consumer no longer strands its messages. Aborting the signal passed to
consume() was checked only between polls. A long poll open at the abort stayed open for up to its
timeout, and the broker could still hand it a message. .each() then dropped that message without
settling it, so its partition stayed blocked until the lease expired, on every rolling restart. The
abort now closes the poll in flight, and the broker hands nothing to a poll whose caller is gone. A
pop answer already arriving is read to the end, since the broker leased its messages when it sent
it. Under .each(), messages popped but not yet handed to the handler, and those popped beyond
.limit(), go back with a retry ack: the lease is released and no retry is charged. An aborted
request is not a backend failure: it is not retried, does not fail over to another node and does not
mark one unhealthy. A wait between attempts (429 backoff, retry after a 5xx or a network error) ends
at once instead of running out. A wait(false) consumer stopped during a pop now resolves instead of
rejecting. Two cases still fall back to the lease: an answer lost before its headers arrive, and a
retry ack that cannot be delivered. Measured on a three-node 2.0.1 cluster, a consumer stopped as a
message arrived: 40 of 40 messages waited out the lease before, none after.
JS client: the request timeout covers the response body. A JSON response was read after the
request’s try block had ended, so the timeout was already cleared: a body that stalled midway hung
with no timeout.
2.0.2 - 2026-10-07
C++ client: commit_on_delivery() commits a pop at delivery. The new
QueueBuilder::commit_on_delivery() makes pop() and pop_result() send the broker’s
autoAck=true: the broker moves the consumer group’s cursor past the messages as it hands them
out, with no lease and nothing to ack, so the delivery is at-most-once. auto_ack() still decides
only the ack after a consume() handler and has no effect on pop(). consume() always leases
its messages, and with commit_on_delivery() it throws std::invalid_argument before any request.
On the ephemeral pop, EphemeralPopOptions::commit_on_delivery is the new name of auto_ack, which
still works and is deprecated.
C++ client: a handler that throws no longer loses a message. With auto_ack(false) and
each(), a handler exception was dropped without a log line and the loop went on with the rest of
the pop: a handler that then acked a later message of the same partition moved the cursor past the
failed one, and the failed message was lost. After a failure, the later messages of the same
partition in that pop are no longer handled; they come back with the failed one. The other
partitions of a multi-partition pop are still handled. The consumer still sends no nack under
auto_ack(false), by design: the failure is now logged, and the message comes back when its lease
expires. With auto_ack(true) the nack now carries the exception’s message as its error, capped at
4096 bytes and with invalid UTF-8 replaced, so the text cannot stop the nack. A handler that throws
something other than a std::exception is handled the same way; before, it ended the worker.
C++ client: renew_lease() renews. The consume loop accepted the setting and never renewed.
While the handler runs, it now renews the batch’s lease every interval_millis, one request per
lease, and stops after the ack or nack, as the JS, Go and Rust clients do.
C++ client: consume() throws a pop’s 4xx and waits out a 5xx. A 4xx on a pop other than a
403 or a 429, such as a 400, ended the worker inside a pool task whose result nobody read, so
consume() returned as if it had finished. It now throws that error once every worker has
stopped, as it does for a 403. A 5xx that outlasted the retries ended the worker the same way; the
loop now waits a second and polls again, as after a network fault, and a connection timeout now
counts as a network fault. pop() still logs a failure and returns an empty result.
C++ client: ack() and renew() report what the broker refused. ack() returned
success: true for any HTTP 200 and renew() for any answer, but the broker answers 200 with
success: false when it settled or extended nothing, for example under an expired or released
lease. Both now read the body: success is false with an error, and ack() keeps the broker’s
answer in result.
Server: the pop parameter commitOnDelivery. commitOnDelivery=true is the new name of a
pop’s autoAck=true: the broker moves the group’s cursor past the messages as it hands them out,
with no lease and nothing to ack, so delivery is at-most-once. Every SDK also has an autoAck() on
consume(), which is client-side: the pop stays leased and the SDK acks after the handler, so
delivery stays at-least-once. The broker option now has a name of its own. autoAck is still
accepted as a deprecated alias, so the released Go and Rust SDKs and the CLI keep working: a pop
that sends either name set to true commits at delivery, and a pop that sends neither is leased as
before. The queue, partition, discovery and ephemeral pops read both names, conflation refuses
both with a 400 that names both, and the OpenAPI document marks autoAck deprecated. The docs
name only commitOnDelivery. A broker up to 2.0.1 reads only autoAck: it ignores
commitOnDelivery and leases the batch.
Server: a page of message history reads only the appends that hold it. A historical
GET /api/v1/messages read each partition backwards from its live tail, 256 offsets at a time, and
decoded every payload before it applied to, so a small page of old messages could read a large
newer suffix, and a smaller limit did not bound the work. It now seeks the queue log’s indexes,
active and sealed, by timestamp and offset, ranks the candidates from their metadata, and reads
only the appends that contain the page. The response fields, the status filters and the minute
rounding of to are unchanged, and so is the storage format. Listing a tenant’s partitions and
cursors still costs what it did, so deep offsets and selective status filters can still be slow.
Dashboard: the Overview lists the queues that need you, and there is a Supervisors page. Under the verdict, the open issues are a list you can search and filter (needs attention, a lag of 5 minutes or more, no reader, pending increased, all queues), ten to a page. Selecting one opens a drawer beside the page with its pending and in-flight counts, the change between two readings, the groups behind and the next thing to check. The new Supervisors page shows the worker status that Laravel supervisors publish to the broker, and reads it again every 30 seconds like the other pages.
Python client: admin.move_message_to_dlq() and admin.clear_queue() are deprecated and
raise. They sent POST /api/v1/messages/:partitionId/:transactionId/dlq and
DELETE /api/v1/queues/:name/clear, routes the 2.x broker does not have, so every call failed
with a 404 no_such_route. Both now raise NotImplementedError before any request and name the
way that works. For a dead letter, ack the message with the dlq status,
await queen.ack(message, 'dlq', {'group': group}), while its consumer group holds the lease. To
skip what is queued, seek the consumer group to the end,
await queen.admin.seek_consumer_group(group, queue, {'toEnd': True}); to drop the queue with its
messages and configuration, await queen.queue(name).delete().
Python client: renew() and the batch ack() read the broker’s verdict. The broker answers
HTTP 200 whether or not it extended a lease or took an ack, and says which in the body. renew()
reported success: True for every 200, and a batch ack() did the same for a refused ack, for
example under an expired lease. renew() now reports success: False, with the broker’s error,
when it extended nothing: the lease expired, was released by an ack or nack, or does not exist. It
also returns renewed, the count the broker extended. A batch ack() reports success: False
when the broker refused any item, with each item’s verdict in results.
Python client: commit_on_delivery() commits a pop at delivery. The broker can move a
consumer group’s cursor past the messages as it hands them out, with no lease and nothing to ack,
but this client could not ask for it: auto_ack(True) on a pop never reached the broker.
queen.queue(q).group(g).commit_on_delivery().pop(), and pop_result(), now send autoAck=true,
the parameter every 2.x broker reads; the messages come back with no leaseId. This is
at-most-once: a crash after the pop loses the messages. consume() always leases, so it raises
ValueError before any request when the builder has commit_on_delivery(). auto_ack() stays the
ack consume() sends after the handler and still never reaches the broker. On ephemeral queues,
queen.ephemeral.pop() takes commit_on_delivery=True; its auto_ack argument, which meant the
same, still works and is deprecated. POP_DEFAULTS now says what a pop does: wait is True, as
every pop has long-polled unless .wait(False), and auto_ack is never sent. No other behaviour
changes.
Python client: pop() raises the broker’s 400 refusal of a conflating pop. pop() returns
an empty list when a pop fails, and did so for the broker’s 400 refusal of a conflating pop too:
one without a consumer group, or one with commit_on_delivery(). No retry can make either
succeed, and an empty list reads as an empty queue, so a consumer that asked for last-value
delivery never learned that it was not getting it. pop() and pop_result() now raise that 400,
as the JavaScript client does. Every other failure still returns an empty list.
Python client: a nack in each() mode skips only its own partition. A multi-partition pop
claims several partitions under one lease, and a nack releases only the failed message’s
partition. The loop dropped the whole rest of the pop after a nack, so the other partitions’
messages stayed leased and came back only when the lease expired. It now skips only the later
messages of the failed partition, which the broker redelivers, and handles the others at once.
Python client: a handler error is never taken for a pop error. With auto_ack(False),
consume() sends no nack for a handler that raises, by design, and the error stops the consumer.
When the error’s text said “timeout” or “connection”, the loop read it as a long-poll timeout or a
network fault instead: it polled again with the message still leased, and the error was lost. Such
an error now stops the consumer like any other handler error.
Laravel and supervisor 0.8.0: prefork per pool. A pool’s own prefork key wins over
supervisor.prefork: false spawns that pool’s workers, true forks them, null follows the
switch. One forking pool is enough to start the fork server, and both engines spawn the workers of
a pool with prefork false beside it. Use it for jobs that must not run in a forked process, such
as a Kafka client the job creates; a boot that starts a thread still makes the fork server refuse to
serve, and then every pool spawns. The dashboard names the pools that differ from the switch. PHP
client 2.3.0 pins supervisor 0.8.0, which reads the new key: run php artisan queen:supervisor-install after the upgrade. A fork took 0.48 ms on the Linux server, against 88 ms
for a worker that booted the benchmark application on its own
(benchmark-queen/2026-10-05-laravel-worker-memory).
PHP client: autoAck() is the ack consume() sends after the handler. The reference said
that on pop() it was the broker’s at-most-once auto-ack, which never reached the broker. It is not
meant to: autoAck() has no effect on pop(), popResult() and popDetached(). A pop commits at
delivery only with the new commitOnDelivery(), below. The reference and the builder now say so,
and tests pin it. No behaviour changes.
PHP client: commitOnDelivery() commits a pop at delivery. The broker can move a consumer
group’s cursor past the messages as it hands them out, with no lease and nothing to ack, but this
client could not ask for it. $queen->queue($q)->group($g)->commitOnDelivery()->pop(), and
popResult() and popDetached(), now send autoAck=true, the parameter every 2.x broker reads,
and nothing else; the messages come back with an empty leaseId. This is at-most-once: a crash
after the pop loses the messages. consume() and getConsumer() always lease their messages, so
they throw LogicException before any request when the builder has commitOnDelivery().
autoAck() stays the ack consume() sends after the handler and still never reaches the broker.
The broker refuses commitOnDelivery() together with conflation() (400), and pop() throws
that HttpException. On ephemeral queues, $queen->ephemeral()->pop() takes
'commitOnDelivery' => true; its autoAck option, which meant the same, still works and is
deprecated, and passing both throws InvalidArgumentException.
PHP client: Admin::moveMessageToDLQ() is deprecated and throws. It posted to
/api/v1/messages/:partitionId/:transactionId/dlq, a route the 2.x broker does not have, so every
call failed with a 404 no_such_route. No 2.x route dead-letters a message by its address. The
method now throws BadMethodCallException before any request and names the way that works: ack
the message with the dlq status, $queen->ack($message, 'dlq', ['group' => $group]), while its
consumer group holds the lease.
PHP client: renewLease() in batch mode counts from the pop. With concurrency(N), consume()
handles the batches of a poll round one after the other, but it started each batch’s renewal
interval when the batch reached its handler, so the check before that handler never found it due.
The interval now runs from the pop answer, and a batch that waited behind other handlers longer
than the interval is renewed before its own handler starts. The loop still cannot renew while a
handler runs, since PHP runs the handler on the loop’s only thread, so a handler that declares a
second parameter now gets a renewal call: function (array $messages, \Closure $renew). $renew()
renews the leases in hand once renewLease()’s interval has passed, or on every call without an
interval, and returns whether the broker extended them; a long handler calls it between parts of
its work. A one-parameter handler is called as before. The loop’s own renewal sends one request per
lease instead of one per message.
Laravel: queen:consume sets and renews its lease, and stops on --idle-timeout. With
--auto-ack, the nack of a handler that throws now carries the exception’s message as its error.
Without --auto-ack the command still sends no nack, as the SDK consumers do: it prints the error
and Not nacked without --auto-ack, and the message comes back when its lease expires.
--lease=SECONDS, default retry_after (90), sets the lease of every pop; with
lease_renewal on in config/queen.php and --auto-ack, the command renews it while handle()
runs, with the renewer queue:work uses, checks it before the ack or nack, and neither acks nor
nacks a lease it can no longer vouch for. Without --auto-ack renewal stays off and the command
says so: a handler that acks by itself releases the lease, and renewing it would stop the process.
--idle-timeout, which did nothing, now stops the command with exit code 0 after N ms without a
message. --limit counts every message handed to handle(), failed ones too, and a pop never asks
for more than it leaves. A refused ack or nack prints one warning, and a broker the pops cannot
reach is reported at most every 30 s, then once when it answers again; the command waits 1 s
after each pop that reached no broker instead of polling in a tight loop. HighLevelConsumer gains
lastPopError(), and its ack() and nack() take an optional error. The help of --batch and
--conflation now says what they do.
Laravel: the lease renewal helper finds the application’s autoloader. It walked up from
the package’s Queen.php to the first vendor/autoload.php. When Composer links the package in
from a path repository, PHP reports that file by the link’s target, outside the application, so
the helper failed to start (“Unable to locate Composer autoload.php”) or loaded the package’s own
development autoloader. It now asks Composer for the autoloader that loaded the package, and
walks up only when Composer registers none.
Laravel: the supervisor checks the settings its workers run with, on every connection. A
worker’s queue connection starts from config/queen.php whatever its name, but the supervisor
and the dashboard read config/queen.php only under the connection named queen. A pool on a
second connection, for example one that sets only 'prefetch' => 'auto', was refused for a
missing lease_renewal that its workers did have, and the dashboard showed lease renewal off for
pools whose workers renewed. Both now start every queen connection from config/queen.php, as
the workers do.
Go client: Each() with AutoAck(false) stops at the first handler error. The loop handed
the rest of the popped batch to the handler after a message failed and kept only the last
message’s error. When a later message succeeded, Execute returned nil, and if the handler had
acked that later message, the ack moved the cursor past the failed one, so it was never delivered
again and never reached the dead-letter queue. Now the first error stops the worker at that
message, the rest of the batch does not reach the handler, and the error comes back out of
Execute, as the transaction tutorial says. The loop still sends no nack under AutoAck(false):
the failed message comes back when its lease expires, which spends no retry, so nack it in the
handler when it should count against the queue’s RetryLimit.
Go client: Renew() reports a lease the broker did not extend. POST /api/v1/lease/:leaseId/extend answers 200 with success: false and renewed: 0 when the lease
expired, was released by an ack or nack, or never existed, and Renew() reported every 200 as
Success: true. It now reads success from the body and fills Error when it is false, as the
JavaScript and Rust clients do.
Go client: a pop the broker refuses no longer loops. The consume loop answered any error
other than a network error, a 403 or a 429 by popping again at once: a token the broker refused
(401) brought 14,546 pops in one second, and a conflating consumer without a group 2,937 pops in
half a second, with nothing shown unless logging was on. A 4xx now stops the worker and comes back
out of Execute, as a 403 does; a 5xx waits a second before the next pop, as a network error
does.
Go client: a nack in Each() mode skips only its own partition. A multi-partition pop claims
several partitions under one lease, and a nack releases only the failed message’s partition. With
AutoAck, the loop dropped the whole rest of the pop after a nack, so the other partitions’
messages stayed leased and came back only when the lease expired: live, with a 30 s lease, partition
B was not handled before the run ended. It now skips only the later messages of the failed
partition, which the broker redelivers, and handles the others at once. With AutoAck(false) a
handler error still stops the consumer at that message.
Go client: CommitOnDelivery() is the pop’s option, and AutoAck() no longer affects a pop.
Behaviour change. On a pop, the broker’s autoAck=true moves the consumer group’s cursor past the
messages as it hands them out: no lease, nothing to ack, at-most-once. AutoAck() sent it from
Pop and PopResult, while on Consume it is the ack the loop sends after the handler, which
never reaches the wire. The pop now has its own option: CommitOnDelivery(true) sends
autoAck=true from Pop and PopResult, and AutoAck() applies to Consume and ConsumeBatch
only. A pop after AutoAck(true) is now leased, so it is at-least-once (a message can come again,
none is lost) and you ack what it returns; call CommitOnDelivery(true) there to keep committing at
delivery. Consume and ConsumeBatch refuse a builder with CommitOnDelivery(true): Execute
returns ErrCommitOnDeliveryConsume before any request. EphemeralPopOptions gains
CommitOnDelivery, and its AutoAck stays as a deprecated alias with the same effect.
CLI: queenctl pop --commit-on-delivery. The broker moves the group’s cursor past the messages
as it hands them out: no lease, nothing to ack, at-most-once. ephemeral pop takes the same flag.
On both, --auto-ack is now a hidden, deprecated alias with the same effect, and using it prints a
deprecation line on stderr. tail --auto-ack keeps its name, and its help now says what it does:
the client acks each message after printing it (the help said “ack server-side”). bench still
drains with pops that commit on delivery. Until a client-go release has CommitOnDelivery, queenctl
calls CommitOnDelivery or AutoAck, whichever the client-go it is built with has, so the
go install build (client-go v2.0.0) and the workspace build both send autoAck=true.
Rust client: a nack in consume() skips only its own partition. A multi-partition pop claims
several partitions under one lease, and a nack releases only the failed message’s partition. With
auto_ack on, the loop dropped the whole rest of the pop after a nack, so the other partitions’
messages stayed leased and came back only when the lease expired: live, with a 30 s lease,
partition B was not handled before the run ended. It now skips only the later messages of the failed
partition, which the broker redelivers, and handles the others at once. With auto_ack(false) the
loop still abandons the rest of the pop after a failure.
Rust client: the consume loop logs a refused ack. The broker refuses an ack, for example one
sent after the lease expired, with HTTP 200 and success: false on the item. The loop read only the
transport result, so a handler that outlived its lease had its ack refused without a trace.
consume() and consume_batch() now read the verdict and log a refused ack or nack at error level,
with the broker’s reason and, for a batch, how many items it refused. The loop still carries on, and
ConsumeSummary still counts what the loop decided, not what the broker accepted. Under
auto_ack(false) the loop sends no nack for a handler that returns Err, by design, but it logged
“nacked” all the same, and consume_batch() logged nothing. Both now log the handler’s error as a
warning that says the message was not nacked.
Rust client: commit_on_delivery() replaces pop_auto_ack(). On a pop, the broker’s
autoAck=true moves the consumer group’s cursor past the messages as it hands them out: no lease,
nothing to ack, at-most-once. The builder now has an option for it, commit_on_delivery(true),
which pop() and pop_result() send; without it both stay leased, as before. pop_auto_ack() is
deprecated in favour of commit_on_delivery(true).pop() and sends the same request. consume() and
consume_batch() refuse a builder with commit_on_delivery(true) with Error::Invalid before any
request, since a consumer always leases its messages; auto_ack() stays the loop’s ack after the
handler and never reaches the wire. The ephemeral pop builder gains commit_on_delivery() too, and
its auto_ack() is a deprecated alias with the same effect.
JavaScript client: a nack in each() mode skips only its own partition. A multi-partition
pop claims several partitions under one lease, and a nack releases only the failed message’s
partition. The loop dropped the whole rest of the pop after a nack, so the other partitions’
messages stayed leased and came back only when the lease expired: live, with a 6 s lease, the two
messages of partition B waited 6 s after a message of partition A failed. It now skips only the
later messages of the failed partition, which the broker redelivers, and handles the others at
once.
JavaScript client: the pop defaults say what pop() does. POP_DEFAULTS and the README said
a pop returns at once and that autoAck(true) commits it at delivery. Neither was ever true: a pop
long-polls for timeoutMillis unless .wait(false), and autoAck() never reaches the broker,
because it is the ack consume() sends after the handler. A pop commits at delivery only with the
new commitOnDelivery(), below. POP_DEFAULTS.wait is now true, the README and the builder’s
comments say so, and tests pin both. No behaviour changes.
JavaScript client: commitOnDelivery() commits a pop at delivery. The broker can move a
consumer group’s cursor past the messages as it hands them out, with no lease and nothing to ack,
but this client could not ask for it: autoAck(true) on a pop never reached the broker.
queen.queue(q).group(g).commitOnDelivery().pop(), and popResult(), now send autoAck=true, the
parameter every 2.x broker reads, and nothing else; the messages come back with an empty leaseId.
This is at-most-once: a crash after the pop loses the messages. consume() always leases, so it
throws before any request when the builder has commitOnDelivery(). autoAck() stays the ack
consume() sends after the handler and still never reaches the broker. The broker refuses
commitOnDelivery() together with conflation() (400), and pop() raises that 400. On ephemeral
queues, queen.ephemeral.pop() takes commitOnDelivery: true; its autoAck option, which meant
the same, still works and is deprecated, and passing both throws.
JavaScript client: admin.moveMessageToDLQ() and admin.clearQueue() are deprecated and
throw. They sent POST /api/v1/messages/:partitionId/:transactionId/dlq and
DELETE /api/v1/queues/:name/clear, routes the 2.x broker does not have, so every call failed
with a 404 no_such_route and the message not found. Both now reject before any request and
name the way that works. For a dead letter, ack the message with the dlq status,
queen.ack(message, 'dlq', { group }), while its consumer group holds the lease. To skip what is
queued, seek each consumer group to the end, queen.admin.seekConsumerGroup(group, queue, { toEnd: true }); a pop without a group reads as the group __QUEUE_MODE__.
PHP client 2.2.0 - 2026-10-06
Laravel: prefetch 'auto'. Each worker sizes its next pop from how long its jobs take, so a
batch holds about 250 ms of work: short jobs get batches of up to 16, a job of a second or more
gets one per pop. A job is timed from the moment the worker hands it to Laravel to its next pop,
its ACK included, so an empty long poll or the worker’s sleep never counts as work. A queue starts
at one job per pop and doubles only after two pops in a row came back full; a short pop sets the
next one to what the queue had, an empty pop to one job, and slower jobs shrink it at once. A
worker that serves several queues sizes each on its own. It runs in the worker, so it needs no
supervisor and applies at once, and like any prefetch above 1 it needs lease_renewal. On the
Linux server, 32 workers ran 2,003 jobs/s of 10 ms jobs with it against 1,548 at prefetch 1, and
2,867 with ack_async and pop_ahead, as many as a fixed prefetch of 4 with a third of its pops
(benchmark-queen/2026-10-05-laravel-auto-prefetch). The Laravel guide now describes three
profiles, safe, balanced and fast; balanced and fast long-poll (block_for 1) with the pool’s
sleep at 0, since workers asleep when a burst began left it to the few awake ones and pushed the
p99 at 500 jobs/s to 338 ms in one run of five. The dashboard shows 'auto', and its advice for
short jobs suggests it.
PHP client 2.1.0, supervisor 0.7.0 - 2026-10-05
PHP client 2.1.0 pins supervisor 0.7.0: the Rust master reads lease_service from its
configuration and handles the queue:restart exit below, so 0.6.0 cannot run with it.
Breaking, Laravel: config/queen.php reads 20 environment variables instead of 95. Only the
values that differ between environments or deployments still read the environment: the broker
URLs and token, the queue, consumer group and partitions, the supervisor’s read token, state
directory and its remote status, prefork and coordination switches, the dashboard’s switch, path,
domain and console URL, the alert mail, the metrics switch and token, and the supervisor binary’s
install path and mirror. Every other setting is a plain value in the published file, with the
default it had: set it there, or add your own env() call where a value must differ per
environment. QUEEN_SUPERVISOR_LEASE_SERVICE, which the Rust master read from its environment,
is the config key supervisor.lease_service (default on), so the dashboard now shows the master’s
value instead of the web host’s. supervisor.remote_status.key is no longer required: it defaults
to a slug of APP_NAME and APP_ENV. An application that publishes its own config/queen.php
keeps every variable that file reads; the configuration reference maps each removed variable to its
key.
Laravel prefork: php artisan queue:restart runs the deployed code. A forked worker stops at
queue:restart like a spawned one, but its replacement was forked from the fork server, which
still held the code it booted, so the workers kept the old code until the master restarted. A
worker that stops for the restart signal now says so (Laravel 12 gives the reason), and both
engines then start a new fork server: the workers forked from then on boot nothing and run the
code on disk. The old server stays open until the last worker it forked exits. Changes to
config/queen.php still need queen:supervisor terminate, as before.
Laravel prefork: a fork server that is not safe to fork is not used. fork() copies only the
calling thread, so a booted application that runs another thread (a gRPC or Kafka extension, an
APM agent) would give every worker a broken copy. On Linux the fork server now refuses to serve
when another thread outlives a 5-second grace (libcurl’s resolver thread ends within it), names
the threads, and the master spawns its workers instead. It warns about sockets the boot left open,
which every forked worker would share, and it releases the database and Redis connections, log
channels and mailers the boot opened before the first fork, not in each child only.
2.0.1 - 2026-10-05
Everything in 2.0.1-beta below, and:
Traces live on disk. A trace was a row in every node’s RAM, about one and a half times its
stored size, until QUEEN_RAFT_TRACE_RETENTION_S (7 days by default) expired it, so a client
that traced every message with a few KB of data grew every node by hundreds of MB an hour. Each
node now keeps its traces in a second LMDB environment, <QUEEN_RAFT_DIR>/traces/, whose pages
are page cache the kernel can drop: 100,000 traces of 5 KB add 56 MB to a node’s process memory
instead of 810 MB. The trace routes, their fields, order and pagination are unchanged, and traces
written before the upgrade stay readable until they expire. Expiry runs in bounded steps (512
traces each, a few ms at most) instead of one step for everything past the cutoff (139,000
traces took 1.24 s on 2026-10-05). New gauges: queen_raft_traces_map_bytes and
queen_raft_traces_stored.
The upgrade is one-way. A cluster writes traces to disk once every member runs 2.0.1: the
leader then raises the cluster version to 4, and from then on a 2.0.1-beta or older binary
refuses to start on the data directory. On a cluster at version 4, a learner can be added only
once its node is running and answering (as QUEEN_RAFT_JOIN already requires).
A trace with a very long transaction id or name no longer stops the cluster. Its store key
went past LMDB’s 511-byte limit and every node stopped applying at that entry. The request is now
refused with a 400 (name_too_long).
The dashboard shows the memory a node holds. A node’s memory meter showed its resident set,
which counts the file pages the process maps: the store’s LMDB file is read whole at boot, so a
node looked about 1 GB fuller than it was. The meter now shows the process’s anonymous memory,
the part that can run a node out of memory, reported as anonBytes in /api/v1/raft/members;
an older broker still shows its resident set.
2.0.1-beta - 2026-10-05
A new partition’s first message reaches every consumer group. On a queue read by two or more consumer groups, a group that took a new partition into the leader’s memory before the partition’s first message was written could miss that message until the partition’s next message or a change of leader. A second group loading the partition a moment later raised the shared tail and armed only its own cursor, so the append’s own wake-up found nothing left to do. A group’s load now arms every group already watching the partition whenever it finds the tail moved.
The leader keeps only the consumer state that is in use. The consumption engine kept every
(consumer group, partition) pair it had served in the leader’s memory, about 1 KB each, until the
partition or the group was deleted, and a partition is deleted only after PARTITION_CLEANUP_DAYS
(30 by default). A workload that keeps opening partitions, one per conversation or entity, grew the
leader without bound. A whole-queue group’s pair that holds nothing (no lease, nothing to deliver,
no hold, every change durable) and stays so for QUEEN_CONSUME_IDLE_UNLOAD_S (default 600; 0
keeps every pair) now leaves memory. The next message on its partition loads it back from its
cursor row, as a new leader does. New gauges: queen_consume_engine{kind="groups"|"parts"|"partitions"}
and queen_consume_parts_total{kind="loaded"|"unloaded"}.
2.0.0 - 2026-10-03
The PostgreSQL storage class is removed. Queen 2.0 has one storage class, its own replicated
log, and there is no database to run beside it. Each node keeps its whole state in one data
directory, QUEEN_RAFT_DIR (default /var/lib/queen/raft): the queue logs that hold the payloads,
an embedded ordered store for everything the broker looks up, and, on a cluster, the raft state.
A single node answers a write once it is fsynced to its own disk. A cluster of three or five voters
(QUEEN_RAFT_REPLICATOR=openraft, QUEEN_RAFT_NODE_ID, QUEEN_RAFT_PEERS, QUEEN_RAFT_TOKEN)
has one leader order every write, answers it once a majority of the voters have it on disk, and
keeps serving through the loss of a minority. QUEEN_STORAGE is gone with the choice it made.
There is no in-place upgrade from 1.x: a 2.0 broker does not read a 1.x database and messages do
not carry over, so a move is a cutover. Start 2.0 beside the old deployment, re-apply the queue
configuration, move the producers, drain the old deployment and move the consumers.
The SQS facade is removed. It is not in the image or the binary any more, and QUEEN_SQS_*
is not read.
The S3 sink runs inside the broker, on every node. QUEEN_S3_EMBEDDED=true starts the
data-lake sink in the broker process, on a runtime of its own (QUEEN_S3_THREADS, default a
quarter of the cores, 1 or 2); there is no queen-s3 binary, no child process, no QUEEN_S3_BIN
and no /healthz or /metrics listener of its own, and QUEEN_S3_LISTEN and
QUEEN_S3_LOG_FORMAT are named in a warning and ignored. Every node runs it with the same
configuration, one sink per broker tenant (below). Per queue, a lease in the key/value store
(s3:<sink>:<queue>:lease, TTL QUEEN_S3_LEASE_TTL_MS, default 30 s, refreshed every third of
it) picks the node that writes the queue, and every window intent and commit carries that lease as
a required conditional write, so a node that lost the queue cannot commit. A node claims a free
queue after waiting one second for every queue it already runs or is claiming, plus a jitter under
200 ms, at most half the lease TTL, so the least-loaded node claims first and the queues spread
over the nodes; a claim that fails for a passing reason (no leader yet) is retried within about a
second rather than a TTL. Queues are then rebalanced: each sink counts the nodes running it through
TTL’d presence rows, and a node holding more than ceil(queues / nodes) gives one queue back at a
time (drained and released as at a SIGTERM) to a node below its share, so pods started one after
another still end up sharing the queues, and nothing moves in a steady state. A node stopped
with SIGTERM gives its queues back at once; one that dies loses them when its leases expire, and
one that comes back under the same QUEEN_S3_INSTANCE (by default node-<id>@<host>) takes its
own back at once. Each node reads its own applied copy of the
log, followers included, and closes a window against its own safeTime; the lease refresh is a
log entry, so safeTime keeps moving on a broker nothing writes to and the last window before a
quiet spell still closes by age. The sink needs no QUEEN_URL
and no token, and the proxy’s key/value carve-out for it is gone. Both record envelopes are the
1.5.0 formats; every key gains a tenant= level, the manifest a tenant field, the Parquet
footer a queen.tenant pair, and the position checkpoint each partition’s id, so a 2.0 sink is
pointed at a new prefix or bucket. A record’s ts is the stamp of the append that wrote it, and across the lake
(partition, offset, ts) is unique, since a partition deleted and created again starts again at
offset 0. A missing or bad value of one of the sink’s variables fails the broker’s boot, naming
the variable; a bucket that does not answer delays the sink and nothing else. The window buffers are the broker’s memory
now, so QUEEN_S3_MEMORY_MB defaults to 512 instead of 1024, and an out-of-memory takes the
broker: size the container for both. On SIGTERM the sink stops reading at once, finishes the
window it is committing and gives its leases back while the broker hands off and drains, and the
broker waits for it up to QUEEN_S3_SHUTDOWN_GRACE_MS (30 s) counted from the signal. Its state is the s3 block of GET /status, and its
queen_s3_* families are on /metrics/prometheus. queen_s3_lag_seconds is the node’s
safeTime minus the stamp the queue’s lake is complete through (completeThrough in the status),
so a queue read to its end and idle lags by about the guard and one discovery interval instead of
growing; a node exports a queue’s gauges only while it runs the queue. The commit pointer’s tEnd is now an
ISO-8601 timestamp: 1.5.0 wrote integer microseconds, which the retention hold could not read,
so retentionSinkHold always sat at its cap; it now follows the sink. Running it is
deploy/s3; what it writes is
reference/s3.
Every broker tenant can have an S3 sink and a bucket of its own. The default tenant’s sink is
configured by QUEEN_S3_* and turned on by QUEEN_S3_QUEUES: without it the default tenant has no
sink, and another per-tenant QUEEN_S3_* variable set without it fails the boot. Every other
tenant’s is set by the control plane: PUT /api/cp/clusters/:slug/s3 with the tenant’s endpoint,
region, bucket, access key and queues, any other per-tenant setting, and secretKey, which is
stored only sealed with the cell’s QUEEN_ENCRYPTION_KEY (the same on every node; a cell without
one refuses the secret). GET answers it without the secret, DELETE removes it, and the
tenant’s purge and delete remove it too. The routes are offered only where the node runs the sink.
Every node reads those rows every 5 seconds and runs a tenant’s sink while its row is enabled and
its cluster’s status, combined with its tenant’s, is active or push_blocked; a write that moves
updatedAt, which every PUT with a secret does, rebuilds the sink once the old one has drained. The memory budget, threads, fetch
concurrency, discovery interval, guard, lease TTL, multipart threshold, checkpoint cadence and
instance name stay node-wide, in the environment, shared by every tenant’s sink, and a tenant’s
settings cannot name them. Keys start with the tenant, <prefix>/tenant=<id>/queue=<name>/…, and
the sidecars sit under <prefix>/_queen/tenant=<id>/queue=<name>/, so two tenants never share a
key even in one bucket and prefix. A tenant’s leases and commit pointers are in its own key/value
store, where its queues’ retention hold reads them. The s3 block of GET /status is
{mode, phase, threads, controlPlane, sinks}, one entry per tenant with its source (env or
cp), and a control-plane tenant’s queen_s3_* series carry a tenant label; the node renders
each family once, and a removed tenant’s series go with it.
POST /api/v1/partitions/changed answers a sound safeTime. It is the greatest record stamp
the answering node has applied: every record stamped at or below it is already readable on that
node, so a time window that ends there is complete. 1.x derived it from PostgreSQL’s open
transactions, with a fixed fallback floor (safeTimeDegraded, now always false), and the 2.0
betas answered the node’s wall clock, which a record planned before the answer and applied after
it could fall below. It is per node, and it moves only when an entry applies, so an idle node
answers the same value until the next write. Pages run in partition creation order with and
without since, under an opaque cursor that means the same in both, and a complete pass reads
each of the queue’s partitions once, where every page used to sort the whole queue; a cursor of
the old shapes (n|…, t|…) is BAD_CURSOR. Every partition carries id, its uuid, which is
new when a partition is deleted and created again under the same name. since is read to the
microsecond, and one that is not a timestamp is a 400, as in 1.x, where the 2.0 betas read it
as absent and listed everything.
The standalone proxy is removed. The proxy runs inside the broker process with
QUEEN_PROXY_EMBEDDED=true, fronting PORT, or its own QUEEN_PROXY_PORT while PORT stays the
internal broker port. Its tenants, clusters, users, API-key hashes, plans and usage live in the
broker’s replicated key/value store under a system tenant, so the proxy database, PXDB_* and the
proxy’s SQL migrations are gone, along with the queen-proxy binary and image. A first tenant can
come from QUEEN_PROXY_BOOTSTRAP_* at boot, and the rest through the control plane under
/api/cp/*, guarded by QUEEN_PROXY_CP_TOKEN.
The Kafka facade runs in-process. QUEEN_KAFKA_EMBEDDED=true starts it inside the broker,
calling the broker’s router without a socket; there is no child process and no QUEEN_KAFKA_BIN.
Committed offsets are Queen consumer-group positions by default (QUEEN_KAFKA_OFFSET_STORE,
positions or kv).
Environment variables removed. Every PostgreSQL variable (PG_*, DB_POOL_SIZE,
PG_USE_SSL, PG_SSL_REJECT_UNAUTHORIZED, PG_SSL_ROOT_CERT), the disk spool (FILE_BUFFER_*),
the broker mesh (QUEEN_MESH_*, QUEEN_UDP_*, QUEEN_SYNC_*), the hot-list, fusion, ack-fusion,
pop-fusion and admission knobs, QUEEN_APPLY_SCHEMA, RETENTION_PARALLELISM, the statistics
refresh intervals, QUEEN_STORAGE, QUEEN_SQS_*, QUEEN_S3_BIN and PXDB_*. Some 1.x names stay
because the 2.0 engine reads them: RETENTION_INTERVAL, RETENTION_BATCH_SIZE,
PARTITION_CLEANUP_DAYS, QUEEN_PARTITION_CLEANUP_ENABLED, METRICS_FLUSH_MS,
QUEEN_SWEEPER_BACKOFF_MIN_MS, QUEEN_SWEEPER_BACKOFF_MAX_MS,
QUEEN_SWEEPER_TRANSIENT_BACKOFF_MS, QUEEN_SWEEPER_MAX_ATTEMPTS, QUEEN_STMT_TIMEOUT_MS,
DEFAULT_TIMEOUT, POP_DEFAULT_TIMEOUT_MS and DEFAULT_SUBSCRIPTION_MODE.
Routes and answers. GET /api/v1/analytics/postgres-stats and the broker-to-broker
/internal/api/* routes are gone, and /metrics no longer carries a database block. A push item
answers queued, duplicate or error: with no spool there is no buffered. When a node cannot
take a write the request fails with 503 and Retry-After, and a full disk answers 507.
POST /api/v1/stats/refresh still answers 200 and does nothing, because counters are kept as
entries apply.
Maintenance mode is removed. GET/POST /api/v1/system/maintenance,
GET/POST /api/v1/system/maintenance/pop and GET /api/v1/status/buffers answer 404. No push
or DLQ replay is refused with a maintenance 503, and no pop answers {"messages":[],"paused":true}.
The embedded Rust API loses Broker::set_push_maintenance. The SDKs drop their maintenance calls
(get/setMaintenanceMode and get/setPopMaintenanceMode in JS and PHP, the same four in Go,
get/set_maintenance_mode and get/set_pop_maintenance_mode in Python, and Admin::maintenance,
set_maintenance, pop_maintenance and set_pop_maintenance in Rust) along with their handling
of a paused pop; queenctl maintenance and the dashboard’s maintenance switches are gone. The kv,
timers and ephemeral kill switches and tenant quotas stay.
The embedded Rust API boots on a data directory. BrokerConfig::new().raft(dir), with
raft_disk_pct(high, low) and stmt_timeout_ms; pg(), pg_use_ssl,
pg_ssl_reject_unauthorized, pool_size, apply_schema, spool_dir, retention,
stats_refresh, system_metrics and log_reports are removed. StartError has only Config, and
Broker::shutdown() returns nothing.
One image, one binary. ghcr.io/queen-mq/queen carries the broker, with the proxy, the Kafka
facade and the S3 sink linked in, the dashboard and queenctl; the queen-kafka, queen-sqs and
queen-s3 binaries and the PostgreSQL client tools are gone. The dashboard is raft-only: the
Postgres stats panel, the database pool and the disk-spool cards are removed, and the replicated
log’s status takes their place.
Every client is 2.0.0. The JavaScript, Python and Rust packages (queen-mq on npm, PyPI and
crates.io), the PHP client, the Go module, the C++ header and queenctl move to 2.0.0, with the
Queen 2 broker. Go users change an import path: a major version above 1 is part of a Go module’s
path, so the client is now github.com/smartpricing/queen/clients/client-go/v2 and the CLI
installs with go install github.com/smartpricing/queen/clients/client-cli/v2/cmd/queenctl@latest.
The tags keep their directory prefix (clients/client-go/v2.0.0, clients/client-cli/v2.0.0), and
the client’s has to be pushed first, since the CLI requires it. queen-protocol stays at 1.3.0.
JavaScript, Python, Go, C++ and PHP clients, queenctl: every ACK of a transaction names its own
lease. The transaction builders sent a message’s lease only in requiredLeases. A Queen 2 broker
fences each ACK with the leaseId its operation carries, and lends it the one in requiredLeases
only when the bundle names a single lease, so in a bundle that acked messages of two leases (two
pops, two queues, a Laravel worker handing back two batches) every ACK went unfenced: it could
complete a message that another consumer had taken after the lease expired, and commit the rest
of the bundle with it. Each ACK operation now carries leaseId, and the broker refuses the whole
transaction once any of those leases is no longer the caller’s. requiredLeases is still sent,
and 1.x brokers already read the operation’s lease first. queenctl tx sent no lease at all: it
dropped the bundle file’s requiredLeases. It now sends each ACK’s own leaseId, lends an ACK
without one the lease requiredLeases names when it names exactly one, and exits 1 before sending
anything when an ACK without a lease sits in a bundle that names several. The Rust client already
named the lease on each ACK. PHP 1.9.0 shipped without this fix, its two-batch hand-back included.
Laravel dashboard: the Jobs page reads with a read-only token. It read the job metrics through
POST /api/v1/kv, which takes read-write access, so with a read_bearer_token that may only read
the page showed the metrics as unavailable on any broker that checks tokens. It now reads
POST /api/v1/resources/kv/list, the read route the console’s KV browser uses (broker 1.6.0 and
later), through the new Admin::listKv(). It falls back to getPrefix when the broker has no such
route (404) and when the credential may not read but may use the KV surface (403), such as a
proxy API key with consume and no read scope.
The partitions sunflower names its seeds. Each queue owns a wedge of the flower, and hovering a
seed lights its queue and shows the queue, the partition, its pending and its lag. The per-partition
figures come from a new read-only route, GET /api/v1/resources/partitions?queue=&limit=: the
partitions holding the most pending (default 610), each with pending, processing and
lagSeconds, the age of the oldest message its slowest reader has not consumed (null when caught
up). The Overview asks for it only while it draws one seed per partition.
The broker in the top bar. For operators, every page’s top bar shows the cell’s CPU, memory,
disk and Raft state, each figure its fullest node’s, over a hairline of how full it is; a click
opens every node. Colour appears only past a line: 90% of the CPUs a node may use, 80% and 90% of
its memory limit, and the node’s own disk write gate. GET /api/v1/raft/status and each member of
GET /api/v1/raft/members gain host, the figures behind it.
The sidebar collapses to a rail of icons (the button at the top left of the top bar, or ⌘\),
remembered per browser. Members and Users share one row: an operator switches between the acting
cluster’s members and every account on the cell from the Members page.
PHP client and Laravel supervisor
Laravel: up to 1,024 stripes per queue. A stripe runs one job at a time, so the 64 stripes a
queue could have capped its ordinary jobs at 64 busy workers. QUEEN_PARTITIONS now takes 1 to
1,024 (64 by default, unchanged), and a pop still asks for at most the 64 partitions the broker
checks out per call: the broker serves the next ready ones in turn. On the Linux benchmark server,
128 workers drained 50,000 jobs at 9,257 jobs/s with 256 stripes against 4,722 with 64, 64 workers
20% faster, and the p95 at 500 jobs/s fell from 23.4 to 16.7 ms, with no measured cost at 50 jobs/s
(benchmark-queen/2026-10-02-laravel-partition-stripes). Event-driven scaling watches up to 1,024
stripes per connection in one fetch; the regular poll finds jobs on the others.
Laravel: a crashed worker’s prefetched jobs keep their attempt. With prefetch above 1 and
lease renewal, a worker that dies without a shutdown (SIGKILL, the kernel’s OOM killer, a PHP
fatal error such as memory_limit) no longer leaves its unstarted jobs to lease expiry, which
charged each one an attempt: with tries = 1 they failed without running, and a job that crashed
its worker every time used up the attempts of the jobs prefetched with it. The worker now journals
the transaction its shutdown would send, and rewrites one small record of it before each ACK or
release. Whatever renews its lease sends that transaction once the worker has exited holding the
lease, while the lease is still the worker’s: the Rust supervisor’s master on Linux, whose journal
sits next to its lease socket, or the worker’s PHP lease helper, whose journal sits in a private
temporary directory. Each unstarted job is completed and copied to its partition with the runs so
far, the job that was running counts its run, and the jobs return at once instead of after
retry_after. Every ACK in it names the lease, so the broker refuses the whole transaction once
the lease is no longer the worker’s. Still charged one attempt: a lost node, a crash that takes
the lease helper with the worker, and a batch popped ahead whose answer the worker had not read.
The dashboard’s Configuration page, queen:supervise and the Rust supervisor (through
queen:supervisor-config) now warn at start about a pool with tries 1 on a connection with
prefetch above 1 or pop_ahead.
Laravel: a job handed back unstarted keeps its attempt. A worker that stopped with a
prefetched tail (--memory, --max-jobs, a deploy, Laravel’s timeout handler), or that could
not track a batch popped ahead, handed the jobs back with a retry ACK. The broker counts the
next pop of such a position as a redelivery, so each hand-back charged an attempt to jobs that
never ran, and with tries = 1 they failed with MaxAttemptsExceeded on their next delivery.
The hand-back is now one transaction that completes each unstarted job and pushes a copy to its
partition with the runs so far, as a release does. A crash is handled by the entry above.
Laravel: a forked child no longer kills its worker. With the PHP lease-renewal helper, a
child forked by a job (Laravel’s fork concurrency driver, pcntl_fork()) stopped the parent’s
helper when it exited, and the parent’s watchdog then SIGKILLed the parent mid-job; the child
was also fenced as soon as a subprocess of its own ended.
PHP client: a detached request is written before the next job runs. With ack_async or
pop_ahead, a request on a new connection (a worker’s first detached request, or one after the
broker closed an idle keep-alive connection) was only started when it was sent: cURL wrote it
when the request was settled, which with prefetch above 1 is after the next job. A hard kill
during that job (shutdown grace exceeded, out of memory, node loss) lost the ACK, and the
acknowledged job ran again when its lease expired. postDetached() and getDetached() now
return once the whole request is written, the connection included; on the Guzzle transport
(behind a proxy, or with QUEEN_SDK_HTTP_TRANSPORT=guzzle) this holds for a request with a
body, the ACK. A request that cannot be written within the 5-second connect timeout throws at
once, and the Laravel queue then acknowledges synchronously. A pop sent ahead, which carries no
body, is waited for at most 250 ms, so a slow or dead backend does not hold up the next job.
PHP client: a process forked by a job exits. libcurl’s resolver threads do not survive
fork(), and recent libcurl keeps them alive for a moment after each name resolution (2 seconds
on the multi handle that carries detached requests). A child forked in that window, by Laravel’s
fork concurrency driver or by pcntl_fork() in a job, waited for them forever when its exit
freed the cURL handles it inherited, so the job hung until its timeout and left the child stuck.
The cURL transport now sets CURLOPT_QUICK_EXIT and shares one DNS cache between its handles,
so detached requests reuse the synchronous handle’s resolution and never start a resolver
thread of their own.
Laravel: job metrics and tag records make one bounded attempt. They are written from every
worker’s job events and on WorkerStopping, and used the queue’s ordinary client, whose
retries held the worker on a slow or rate-limiting broker and could spend the shutdown grace
before the prefetched tail was handed back. They now use one 2-second attempt.
Laravel dashboard: retry a failed job in one click. The failed-job page has a Retry now
button. It runs queue:retry for that job, so the broker’s dead-letter entry and the
failed_jobs row stay in step, and the page still shows the command for a terminal. Forgetting,
flushing and pruning stay with Laravel’s commands. An exception during the retry is reported to
the application’s log, and the page shows only its class: its message can quote the job’s
payload. The id all is refused, since queue:retry all retries every failed job.
Laravel dashboard: what each queue holds now. The Workload page shows, for every supervised
queue, the jobs waiting and running and how long the oldest unfinished job has waited, from one
broker read per queue, cached for 5 seconds. QUEEN_DASHBOARD_CONSOLE_URL links each queue to
the Queen console, which lists the messages themselves. An invalid value is reported to the
application’s log and turns the links off; it never stops the application or its workers.
Laravel dashboard: the Configuration page is a tuning guide. It shows every resolved setting
of the connection and the supervisor with its environment variable, each pool as the running
supervisor published it, and advice from the supervisor’s state and the job metrics: job classes
whose longest run exceeds shutdown_grace, one worker for several queues, short jobs on
prefetch 1, prefork off, a lease helper per worker, polling instead of event-driven scaling,
and pools that could run a job twice. A setting that can hold a credential shows only whether it
is set. Job metrics now record each class’s longest run.
Laravel: a rejected broker URL no longer leaks its password. The supervisor configuration printed an invalid broker URL, credentials included, into logs and error pages; it now redacts them.
Supervisors (both engines). A worker that ran at least stable_after and exits non-zero
(queue:work exits 12 at --memory, a job timeout kills the worker) is restarted at once: it
no longer holds its pool at a single probe for stable_after. On stop, the workers get SIGTERM
before the coordination leave and the remote status publish, which a slow broker could stretch
past the platform’s stop deadline. The PHP engine’s event-driven watcher can no longer freeze
the master loop on an answer that stalls mid-body, its fork server survives a failed
pcntl_fork(), and preforked workers honour --quiet. The Rust engine trusts the platform’s
CA store (and SSL_CERT_FILE), so a broker behind a private CA works, and its lease service
fences a worker before it logs why, so a broken stderr pipe cannot skip the fence.
Supervisors: a job timeout is not a crash. Laravel SIGKILLs a worker whose job outlives its
timeout, and both engines counted that exit as a crash: with job timeouts shorter than
stable_after, a burst of them opened the restart circuit and left the pool at one probe worker
for every job on its queue, the healthy ones included. The worker now leaves a marker in the state directory’s exits/ before it dies, and
the master restarts it without backoff and without counting it. A worker that stops at --memory
after it handled a job counts as a clean exit too; one that stops before any job still backs off,
since its boot alone passes the limit. Every other non-zero exit counts as before, a SIGKILL
without a marker included (the OOM killer, a lease fence, kill -9).