Skip to content

Queen MQ changelog: releases since 2.0

Every release since 2.0.0, which replaced PostgreSQL with a Raft log on local disk: the broker, the clients, the Laravel driver and the supervisor, newest first.

Updated View as Markdown

Before 2.0.0 (2026-10-03), Queen MQ kept its data in PostgreSQL. 2.0 replaced that with its own log, replicated with Raft, on each node’s local disk, and a 2.0 node does not read a 1.x database. This page lists every release since 2.0.0, newest first. The 1.x entries are in CHANGELOG.md on GitHub.

2.1.0

Server: a standby cluster. A second cluster can now replay the first one’s log and take over when the first is lost. The standby’s leader reads the source’s committed entries over the raft port and proposes each one into its own log, so the standby holds the source’s messages, cursors, leases, KV, timers and dedup window, a moment behind. It answers reads, refuses writes with 503 and "code": "standby" on every node, and becomes an ordinary cluster with POST /api/v1/system/link/promote. An empty standby follows a young source from the start of its log (QUEEN_LINK_STANDBY); a source with a history seeds the standby’s first node with its snapshot (QUEEN_LINK_SEED). Replication is asynchronous: a promotion after a crash loses at most what the standby had not read yet, and a planned switch loses nothing. On one laptop, with both clusters and the load generator sharing a disk, a standby stayed within 0.25 s of a source taking 200,000 messages a second. The source keeps its log for the standby on every node and across restarts, and gives it up when its own disk fills (QUEEN_LINK_HOLD_S, QUEEN_LINK_HOLD_DISK_PCT). GET /api/v1/system/link, the link block of /health and the queen_link_* series say where a standby is. The source needs QUEEN_LINK_TOKEN, the standby QUEEN_LINK_SOURCE and QUEEN_LINK_SOURCE_TOKEN; a cluster with none of them behaves as before. See Standby cluster.

Server: the last node of a cluster to stop no longer hangs. A node that led a cluster whose other nodes were already down, with an entry in its log that could no longer commit, waited for that entry for ever while it shut down, and used a whole core doing it. A stopping node now leaves once every caller of its entries in flight has an answer; an entry that timed out had already answered retry.

Server: a leader that no majority has acknowledged for two seconds no longer stops for good. A leader cut off for two or three seconds and not replaced, or one whose majority needs a follower with a slow disk, had its next write refused by raft, took the refusal for a lost leadership and stopped planning. It still led, so nothing started it again: every request ran into its deadline until the leadership changed or the node restarted. Every release since 2.0.0 has it. Such a leader now keeps its entries and logs them, in order, when a majority answers again, and a leader that finds itself stopped while it leads starts again after 5 s. Found by Jepsen on a five-node cluster with two slow disks. With the leader of five nodes cut off for 2.4 s, the build before the fix took no write afterwards in 11 of the 12 trials where the leader kept its leadership, and this one took the next write in all 10. The log lines that say when it happens are on Monitoring.

Server: locks and semaphores. A lock is a lease: one holder at a time, for a lifetime the holder declares and renews, with a token that fences a holder that outlived it. A semaphore is the same lease with up to 1,024 permits. One route, POST /api/v1/locks, carries acquire, renew, release and get, up to 64 locks a call, and answers a verdict per operation with HTTP 200 (acquired: false, reason: "held" is not an error). A permit is one KV row in the namespace queen-locks (key <name>#<slot>, the owner as its value, the lifetime as its TTL), so the route adds nothing to the log’s format: an acquire is a putIfAbsent, a renew a put with expect, a release a delete with expect. The token is the row’s version and changes at every renew. An owner makes a call safe to send again: the same owner is answered the permit it has. Locks count against the tenant’s KV quota and KV write rate, need a read-write token, and behind the proxy are part of the KV plan and never quota-blocked. New metrics: queen_locks_ops_total and queen_locks_op_duration_milliseconds. A lock expires; a holder that crashed keeps it until its lifetime ends, a waiter polls, and there are no read/write locks.

Server: KV check, a precondition that writes nothing. {"op":"check","ns","key","expect"} asserts that a key is at a version, or absent with expect: 0. With required: true it gates a KV call or a transaction on a key the call does not write. It is what guards a transaction with a lock: every acquire and renew answers a guard, a check of the permit’s row at its token, and a transaction that carries it commits only while the lock is still the caller’s.

Server: KV versions are a fencing token. On one key, a later write always has a higher version than every earlier one, across a delete and a re-create, an expiry, a restart and a change of leader. The planner always assigned them that way; it is now the documented contract, with a test, and the docs no longer say not to rely on their order.

Clients: lock, semaphore, check and guard, in all six SDKs. queen.lock(name) and queen.semaphore(name, limit), each with a lifetime, return a handle that acquires (with an optional wait), keeps the current token, renews and releases; transaction().guard(lock) puts the guard in the commit and reads the token when the commit is sent. The JavaScript, Python, Go and Rust handles renew in the background every third of the lifetime and signal a lost lock. The PHP and C++ handles have no background renewal: keepAlive() (keep_alive() in C++) renews at a checkpoint in the work loop. A guarded commit that lost only to its own handle’s renewal is sent again with the new token. kv.check is on the KV client and the transaction’s KV builder of each SDK. In the JavaScript, Python, Go, Rust and C++ clients 2.1.0 and the PHP client 2.4.0.

Dashboard: a Locks page. Every held lock and semaphore permit with its holder, since when it is held, its last renewal and when it expires. It reads the permits as KV rows, so a viewer can open it, and it writes nothing: the drawer shows the guard a transaction carries and the call that releases the lease period on screen. get answers the same since for each holder: a renewal does not move it.

Server: a KV call that mixed a read with writes that all lost no longer hangs. A batch on POST /api/v1/kv such as a putIfAbsent that lost beside a get wrote nothing, so it had no log position, and its read waited for one until the statement timeout (30 s) and answered 503. Such a call now takes a position of its own and answers at once.

Server: the upgrade is one-way. A node forwards a check to the leader only when every member can read it. Once every member runs this release the leader raises the cluster version to 5, and from then on 2.0.4 or older refuses to start on the data directory. Until then a call that carries a check answers 503 (kv_check_needs_cluster_version_5); acquire, renew and release need no new format and work as soon as the node you call runs this release. To keep the way back open while the release bakes, set QUEEN_RAFT_CLUSTER_VERSION_MS=0.

2.0.4 - 2026-10-08

Dashboard, sign-in page and docs: the logo is a yellow sunflower. A new drawing replaces the red one of 2.0.3: seven yellow petals round the q of “queen”, with “message queue” under the name. The dashboard shows the q and its petals beside the word “queen”, the sign-in page and the README show the whole logo, and the tab icon is the q with its petals. The picture of the bee and the sunflower on the documentation site takes the same yellow.

Dashboard: the logo’s yellow is the accent. Primary buttons are yellow with black labels, and the page you are on, selected segments and focus rings take the same colour. On light surfaces, text and rings in that colour are a darker gold, so they can be read. The sign-in page and the documentation site follow. Status colours keep their meanings: amber is attention, coral is failure.

Dashboard: System, Light and Dark. The theme toggle in the top bar is now a choice of three. System follows the device’s colour scheme as it changes, and is what a browser with no stored choice gets. Light and Dark stay fixed on that device until System is chosen again, and a choice made before 2.0.4 is kept.

Dashboard: supervisor cards share one grid. The Supervisors page gave each publication group its own two-column grid, so two groups with one instance each took two rows and left the second column empty. The instances now share one grid, each card under its group’s name: two to a row on a wide screen, one on a narrow one.

The broker’s engine, its API and the clients are as in 2.0.3.

2.0.3 - 2026-10-08

Server: a pop that nobody receives no longer counts a delivery attempt. A pop’s answer waits for the checkpoint that holds its lease. When that takes longer than the margin the broker keeps before the request’s deadline, or the client leaves, nobody receives the answer, and the broker hands the leases back at once. It left the delivery attempt counted all the same, so the next consumer got the message as a redelivery. With deliveryAttempt 2 on its first real delivery, a Laravel job with tries = 1 failed with “attempted too many times” without ever running. In a 23-hour run of 4.1 million jobs at 50 jobs/s, 4 jobs failed that way, each about 0.4 s after its push, while every message went out to a worker once; checkpoint writes reached 0.45 s against a 250 ms margin. Handing back an unreceived claim now also takes back its attempt, on the leader’s engine and through the Nack a follower sends. A lease that really expired still counts.

Clients: a consumer can report to the dashboard, off by default. Every SDK’s consumer builder takes a supervision option that names a group for the dashboard: .supervision({ group }) in JavaScript, Supervision(&queen.SupervisionConfig{...}) in Go, supervision(...) in the Python, Rust, C++ and PHP clients. Each consume call then writes an observation to the broker’s queen-supervisor KV namespace every 10 seconds, and a last one when it stops: its execution model, its live loops against the configured concurrency, the handlers running, the handler calls completed and failed, and the age of the oldest handler still running. No payload and no error text is published. Publishing is best effort and bounded, and it needs a credential that can write KV. It only observes: it restarts nothing, scales nothing and changes no ack or lease, and with the option off there is no timer and no request. The PHP client publishes at its cooperative checkpoints, so a long synchronous handler leaves its last observation stale. In the JavaScript client 2.0.4, the Python, Go, Rust and C++ clients 2.0.3 and the PHP client 2.3.1.

CLI: queenctl tail --supervision-group. tail reports the same way when the flag names a group, one identity for each invocation.

Dashboard: the Supervisors page shows the consumers that report. Beside the process supervisors it lists each reporting consumer with its execution model (async tasks, goroutines, threads or cooperative loops), its live loops against the configured concurrency, its busy handlers, its completed and failed handler calls, its last completion and the age of its oldest handler in flight. A completed handler does not prove its ack succeeded, and a fresh heartbeat does not prove progress; a stale or inconsistent report stays unconfirmed. Process budgets and readiness stay with the process supervisors.

Dashboard, sign-in page and docs: a new logo. The q that is a sunflower replaces the sunflower and the bee: as a badge beside the word “queen” in the dashboard, as the wordmark on the sign-in page and in the README, and as the tab icon. The sign-in page now takes the dashboard’s light or dark scheme.

JS client 2.0.3 - 2026-10-07

JS client: stopping a consumer no longer strands its messages. Aborting the signal passed to consume() was checked only between polls. A long poll open at the abort stayed open for up to its timeout, and the broker could still hand it a message. .each() then dropped that message without settling it, so its partition stayed blocked until the lease expired, on every rolling restart. The abort now closes the poll in flight, and the broker hands nothing to a poll whose caller is gone. A pop answer already arriving is read to the end, since the broker leased its messages when it sent it. Under .each(), messages popped but not yet handed to the handler, and those popped beyond .limit(), go back with a retry ack: the lease is released and no retry is charged. An aborted request is not a backend failure: it is not retried, does not fail over to another node and does not mark one unhealthy. A wait between attempts (429 backoff, retry after a 5xx or a network error) ends at once instead of running out. A wait(false) consumer stopped during a pop now resolves instead of rejecting. Two cases still fall back to the lease: an answer lost before its headers arrive, and a retry ack that cannot be delivered. Measured on a three-node 2.0.1 cluster, a consumer stopped as a message arrived: 40 of 40 messages waited out the lease before, none after.

JS client: the request timeout covers the response body. A JSON response was read after the request’s try block had ended, so the timeout was already cleared: a body that stalled midway hung with no timeout.

2.0.2 - 2026-10-07

C++ client: commit_on_delivery() commits a pop at delivery. The new QueueBuilder::commit_on_delivery() makes pop() and pop_result() send the broker’s autoAck=true: the broker moves the consumer group’s cursor past the messages as it hands them out, with no lease and nothing to ack, so the delivery is at-most-once. auto_ack() still decides only the ack after a consume() handler and has no effect on pop(). consume() always leases its messages, and with commit_on_delivery() it throws std::invalid_argument before any request. On the ephemeral pop, EphemeralPopOptions::commit_on_delivery is the new name of auto_ack, which still works and is deprecated.

C++ client: a handler that throws no longer loses a message. With auto_ack(false) and each(), a handler exception was dropped without a log line and the loop went on with the rest of the pop: a handler that then acked a later message of the same partition moved the cursor past the failed one, and the failed message was lost. After a failure, the later messages of the same partition in that pop are no longer handled; they come back with the failed one. The other partitions of a multi-partition pop are still handled. The consumer still sends no nack under auto_ack(false), by design: the failure is now logged, and the message comes back when its lease expires. With auto_ack(true) the nack now carries the exception’s message as its error, capped at 4096 bytes and with invalid UTF-8 replaced, so the text cannot stop the nack. A handler that throws something other than a std::exception is handled the same way; before, it ended the worker.

C++ client: renew_lease() renews. The consume loop accepted the setting and never renewed. While the handler runs, it now renews the batch’s lease every interval_millis, one request per lease, and stops after the ack or nack, as the JS, Go and Rust clients do.

C++ client: consume() throws a pop’s 4xx and waits out a 5xx. A 4xx on a pop other than a 403 or a 429, such as a 400, ended the worker inside a pool task whose result nobody read, so consume() returned as if it had finished. It now throws that error once every worker has stopped, as it does for a 403. A 5xx that outlasted the retries ended the worker the same way; the loop now waits a second and polls again, as after a network fault, and a connection timeout now counts as a network fault. pop() still logs a failure and returns an empty result.

C++ client: ack() and renew() report what the broker refused. ack() returned success: true for any HTTP 200 and renew() for any answer, but the broker answers 200 with success: false when it settled or extended nothing, for example under an expired or released lease. Both now read the body: success is false with an error, and ack() keeps the broker’s answer in result.

Server: the pop parameter commitOnDelivery. commitOnDelivery=true is the new name of a pop’s autoAck=true: the broker moves the group’s cursor past the messages as it hands them out, with no lease and nothing to ack, so delivery is at-most-once. Every SDK also has an autoAck() on consume(), which is client-side: the pop stays leased and the SDK acks after the handler, so delivery stays at-least-once. The broker option now has a name of its own. autoAck is still accepted as a deprecated alias, so the released Go and Rust SDKs and the CLI keep working: a pop that sends either name set to true commits at delivery, and a pop that sends neither is leased as before. The queue, partition, discovery and ephemeral pops read both names, conflation refuses both with a 400 that names both, and the OpenAPI document marks autoAck deprecated. The docs name only commitOnDelivery. A broker up to 2.0.1 reads only autoAck: it ignores commitOnDelivery and leases the batch.

Server: a page of message history reads only the appends that hold it. A historical GET /api/v1/messages read each partition backwards from its live tail, 256 offsets at a time, and decoded every payload before it applied to, so a small page of old messages could read a large newer suffix, and a smaller limit did not bound the work. It now seeks the queue log’s indexes, active and sealed, by timestamp and offset, ranks the candidates from their metadata, and reads only the appends that contain the page. The response fields, the status filters and the minute rounding of to are unchanged, and so is the storage format. Listing a tenant’s partitions and cursors still costs what it did, so deep offsets and selective status filters can still be slow.

Dashboard: the Overview lists the queues that need you, and there is a Supervisors page. Under the verdict, the open issues are a list you can search and filter (needs attention, a lag of 5 minutes or more, no reader, pending increased, all queues), ten to a page. Selecting one opens a drawer beside the page with its pending and in-flight counts, the change between two readings, the groups behind and the next thing to check. The new Supervisors page shows the worker status that Laravel supervisors publish to the broker, and reads it again every 30 seconds like the other pages.

Python client: admin.move_message_to_dlq() and admin.clear_queue() are deprecated and raise. They sent POST /api/v1/messages/:partitionId/:transactionId/dlq and DELETE /api/v1/queues/:name/clear, routes the 2.x broker does not have, so every call failed with a 404 no_such_route. Both now raise NotImplementedError before any request and name the way that works. For a dead letter, ack the message with the dlq status, await queen.ack(message, 'dlq', {'group': group}), while its consumer group holds the lease. To skip what is queued, seek the consumer group to the end, await queen.admin.seek_consumer_group(group, queue, {'toEnd': True}); to drop the queue with its messages and configuration, await queen.queue(name).delete().

Python client: renew() and the batch ack() read the broker’s verdict. The broker answers HTTP 200 whether or not it extended a lease or took an ack, and says which in the body. renew() reported success: True for every 200, and a batch ack() did the same for a refused ack, for example under an expired lease. renew() now reports success: False, with the broker’s error, when it extended nothing: the lease expired, was released by an ack or nack, or does not exist. It also returns renewed, the count the broker extended. A batch ack() reports success: False when the broker refused any item, with each item’s verdict in results.

Python client: commit_on_delivery() commits a pop at delivery. The broker can move a consumer group’s cursor past the messages as it hands them out, with no lease and nothing to ack, but this client could not ask for it: auto_ack(True) on a pop never reached the broker. queen.queue(q).group(g).commit_on_delivery().pop(), and pop_result(), now send autoAck=true, the parameter every 2.x broker reads; the messages come back with no leaseId. This is at-most-once: a crash after the pop loses the messages. consume() always leases, so it raises ValueError before any request when the builder has commit_on_delivery(). auto_ack() stays the ack consume() sends after the handler and still never reaches the broker. On ephemeral queues, queen.ephemeral.pop() takes commit_on_delivery=True; its auto_ack argument, which meant the same, still works and is deprecated. POP_DEFAULTS now says what a pop does: wait is True, as every pop has long-polled unless .wait(False), and auto_ack is never sent. No other behaviour changes.

Python client: pop() raises the broker’s 400 refusal of a conflating pop. pop() returns an empty list when a pop fails, and did so for the broker’s 400 refusal of a conflating pop too: one without a consumer group, or one with commit_on_delivery(). No retry can make either succeed, and an empty list reads as an empty queue, so a consumer that asked for last-value delivery never learned that it was not getting it. pop() and pop_result() now raise that 400, as the JavaScript client does. Every other failure still returns an empty list.

Python client: a nack in each() mode skips only its own partition. A multi-partition pop claims several partitions under one lease, and a nack releases only the failed message’s partition. The loop dropped the whole rest of the pop after a nack, so the other partitions’ messages stayed leased and came back only when the lease expired. It now skips only the later messages of the failed partition, which the broker redelivers, and handles the others at once.

Python client: a handler error is never taken for a pop error. With auto_ack(False), consume() sends no nack for a handler that raises, by design, and the error stops the consumer. When the error’s text said “timeout” or “connection”, the loop read it as a long-poll timeout or a network fault instead: it polled again with the message still leased, and the error was lost. Such an error now stops the consumer like any other handler error.

Laravel and supervisor 0.8.0: prefork per pool. A pool’s own prefork key wins over supervisor.prefork: false spawns that pool’s workers, true forks them, null follows the switch. One forking pool is enough to start the fork server, and both engines spawn the workers of a pool with prefork false beside it. Use it for jobs that must not run in a forked process, such as a Kafka client the job creates; a boot that starts a thread still makes the fork server refuse to serve, and then every pool spawns. The dashboard names the pools that differ from the switch. PHP client 2.3.0 pins supervisor 0.8.0, which reads the new key: run php artisan queen:supervisor-install after the upgrade. A fork took 0.48 ms on the Linux server, against 88 ms for a worker that booted the benchmark application on its own (benchmark-queen/2026-10-05-laravel-worker-memory).

PHP client: autoAck() is the ack consume() sends after the handler. The reference said that on pop() it was the broker’s at-most-once auto-ack, which never reached the broker. It is not meant to: autoAck() has no effect on pop(), popResult() and popDetached(). A pop commits at delivery only with the new commitOnDelivery(), below. The reference and the builder now say so, and tests pin it. No behaviour changes.

PHP client: commitOnDelivery() commits a pop at delivery. The broker can move a consumer group’s cursor past the messages as it hands them out, with no lease and nothing to ack, but this client could not ask for it. $queen->queue($q)->group($g)->commitOnDelivery()->pop(), and popResult() and popDetached(), now send autoAck=true, the parameter every 2.x broker reads, and nothing else; the messages come back with an empty leaseId. This is at-most-once: a crash after the pop loses the messages. consume() and getConsumer() always lease their messages, so they throw LogicException before any request when the builder has commitOnDelivery(). autoAck() stays the ack consume() sends after the handler and still never reaches the broker. The broker refuses commitOnDelivery() together with conflation() (400), and pop() throws that HttpException. On ephemeral queues, $queen->ephemeral()->pop() takes 'commitOnDelivery' => true; its autoAck option, which meant the same, still works and is deprecated, and passing both throws InvalidArgumentException.

PHP client: Admin::moveMessageToDLQ() is deprecated and throws. It posted to /api/v1/messages/:partitionId/:transactionId/dlq, a route the 2.x broker does not have, so every call failed with a 404 no_such_route. No 2.x route dead-letters a message by its address. The method now throws BadMethodCallException before any request and names the way that works: ack the message with the dlq status, $queen->ack($message, 'dlq', ['group' => $group]), while its consumer group holds the lease.

PHP client: renewLease() in batch mode counts from the pop. With concurrency(N), consume() handles the batches of a poll round one after the other, but it started each batch’s renewal interval when the batch reached its handler, so the check before that handler never found it due. The interval now runs from the pop answer, and a batch that waited behind other handlers longer than the interval is renewed before its own handler starts. The loop still cannot renew while a handler runs, since PHP runs the handler on the loop’s only thread, so a handler that declares a second parameter now gets a renewal call: function (array $messages, \Closure $renew). $renew() renews the leases in hand once renewLease()’s interval has passed, or on every call without an interval, and returns whether the broker extended them; a long handler calls it between parts of its work. A one-parameter handler is called as before. The loop’s own renewal sends one request per lease instead of one per message.

Laravel: queen:consume sets and renews its lease, and stops on --idle-timeout. With --auto-ack, the nack of a handler that throws now carries the exception’s message as its error. Without --auto-ack the command still sends no nack, as the SDK consumers do: it prints the error and Not nacked without --auto-ack, and the message comes back when its lease expires. --lease=SECONDS, default retry_after (90), sets the lease of every pop; with lease_renewal on in config/queen.php and --auto-ack, the command renews it while handle() runs, with the renewer queue:work uses, checks it before the ack or nack, and neither acks nor nacks a lease it can no longer vouch for. Without --auto-ack renewal stays off and the command says so: a handler that acks by itself releases the lease, and renewing it would stop the process. --idle-timeout, which did nothing, now stops the command with exit code 0 after N ms without a message. --limit counts every message handed to handle(), failed ones too, and a pop never asks for more than it leaves. A refused ack or nack prints one warning, and a broker the pops cannot reach is reported at most every 30 s, then once when it answers again; the command waits 1 s after each pop that reached no broker instead of polling in a tight loop. HighLevelConsumer gains lastPopError(), and its ack() and nack() take an optional error. The help of --batch and --conflation now says what they do.

Laravel: the lease renewal helper finds the application’s autoloader. It walked up from the package’s Queen.php to the first vendor/autoload.php. When Composer links the package in from a path repository, PHP reports that file by the link’s target, outside the application, so the helper failed to start (“Unable to locate Composer autoload.php”) or loaded the package’s own development autoloader. It now asks Composer for the autoloader that loaded the package, and walks up only when Composer registers none.

Laravel: the supervisor checks the settings its workers run with, on every connection. A worker’s queue connection starts from config/queen.php whatever its name, but the supervisor and the dashboard read config/queen.php only under the connection named queen. A pool on a second connection, for example one that sets only 'prefetch' => 'auto', was refused for a missing lease_renewal that its workers did have, and the dashboard showed lease renewal off for pools whose workers renewed. Both now start every queen connection from config/queen.php, as the workers do.

Go client: Each() with AutoAck(false) stops at the first handler error. The loop handed the rest of the popped batch to the handler after a message failed and kept only the last message’s error. When a later message succeeded, Execute returned nil, and if the handler had acked that later message, the ack moved the cursor past the failed one, so it was never delivered again and never reached the dead-letter queue. Now the first error stops the worker at that message, the rest of the batch does not reach the handler, and the error comes back out of Execute, as the transaction tutorial says. The loop still sends no nack under AutoAck(false): the failed message comes back when its lease expires, which spends no retry, so nack it in the handler when it should count against the queue’s RetryLimit.

Go client: Renew() reports a lease the broker did not extend. POST /api/v1/lease/:leaseId/extend answers 200 with success: false and renewed: 0 when the lease expired, was released by an ack or nack, or never existed, and Renew() reported every 200 as Success: true. It now reads success from the body and fills Error when it is false, as the JavaScript and Rust clients do.

Go client: a pop the broker refuses no longer loops. The consume loop answered any error other than a network error, a 403 or a 429 by popping again at once: a token the broker refused (401) brought 14,546 pops in one second, and a conflating consumer without a group 2,937 pops in half a second, with nothing shown unless logging was on. A 4xx now stops the worker and comes back out of Execute, as a 403 does; a 5xx waits a second before the next pop, as a network error does.

Go client: a nack in Each() mode skips only its own partition. A multi-partition pop claims several partitions under one lease, and a nack releases only the failed message’s partition. With AutoAck, the loop dropped the whole rest of the pop after a nack, so the other partitions’ messages stayed leased and came back only when the lease expired: live, with a 30 s lease, partition B was not handled before the run ended. It now skips only the later messages of the failed partition, which the broker redelivers, and handles the others at once. With AutoAck(false) a handler error still stops the consumer at that message.

Go client: CommitOnDelivery() is the pop’s option, and AutoAck() no longer affects a pop. Behaviour change. On a pop, the broker’s autoAck=true moves the consumer group’s cursor past the messages as it hands them out: no lease, nothing to ack, at-most-once. AutoAck() sent it from Pop and PopResult, while on Consume it is the ack the loop sends after the handler, which never reaches the wire. The pop now has its own option: CommitOnDelivery(true) sends autoAck=true from Pop and PopResult, and AutoAck() applies to Consume and ConsumeBatch only. A pop after AutoAck(true) is now leased, so it is at-least-once (a message can come again, none is lost) and you ack what it returns; call CommitOnDelivery(true) there to keep committing at delivery. Consume and ConsumeBatch refuse a builder with CommitOnDelivery(true): Execute returns ErrCommitOnDeliveryConsume before any request. EphemeralPopOptions gains CommitOnDelivery, and its AutoAck stays as a deprecated alias with the same effect.

CLI: queenctl pop --commit-on-delivery. The broker moves the group’s cursor past the messages as it hands them out: no lease, nothing to ack, at-most-once. ephemeral pop takes the same flag. On both, --auto-ack is now a hidden, deprecated alias with the same effect, and using it prints a deprecation line on stderr. tail --auto-ack keeps its name, and its help now says what it does: the client acks each message after printing it (the help said “ack server-side”). bench still drains with pops that commit on delivery. Until a client-go release has CommitOnDelivery, queenctl calls CommitOnDelivery or AutoAck, whichever the client-go it is built with has, so the go install build (client-go v2.0.0) and the workspace build both send autoAck=true.

Rust client: a nack in consume() skips only its own partition. A multi-partition pop claims several partitions under one lease, and a nack releases only the failed message’s partition. With auto_ack on, the loop dropped the whole rest of the pop after a nack, so the other partitions’ messages stayed leased and came back only when the lease expired: live, with a 30 s lease, partition B was not handled before the run ended. It now skips only the later messages of the failed partition, which the broker redelivers, and handles the others at once. With auto_ack(false) the loop still abandons the rest of the pop after a failure.

Rust client: the consume loop logs a refused ack. The broker refuses an ack, for example one sent after the lease expired, with HTTP 200 and success: false on the item. The loop read only the transport result, so a handler that outlived its lease had its ack refused without a trace. consume() and consume_batch() now read the verdict and log a refused ack or nack at error level, with the broker’s reason and, for a batch, how many items it refused. The loop still carries on, and ConsumeSummary still counts what the loop decided, not what the broker accepted. Under auto_ack(false) the loop sends no nack for a handler that returns Err, by design, but it logged “nacked” all the same, and consume_batch() logged nothing. Both now log the handler’s error as a warning that says the message was not nacked.

Rust client: commit_on_delivery() replaces pop_auto_ack(). On a pop, the broker’s autoAck=true moves the consumer group’s cursor past the messages as it hands them out: no lease, nothing to ack, at-most-once. The builder now has an option for it, commit_on_delivery(true), which pop() and pop_result() send; without it both stay leased, as before. pop_auto_ack() is deprecated in favour of commit_on_delivery(true).pop() and sends the same request. consume() and consume_batch() refuse a builder with commit_on_delivery(true) with Error::Invalid before any request, since a consumer always leases its messages; auto_ack() stays the loop’s ack after the handler and never reaches the wire. The ephemeral pop builder gains commit_on_delivery() too, and its auto_ack() is a deprecated alias with the same effect.

JavaScript client: a nack in each() mode skips only its own partition. A multi-partition pop claims several partitions under one lease, and a nack releases only the failed message’s partition. The loop dropped the whole rest of the pop after a nack, so the other partitions’ messages stayed leased and came back only when the lease expired: live, with a 6 s lease, the two messages of partition B waited 6 s after a message of partition A failed. It now skips only the later messages of the failed partition, which the broker redelivers, and handles the others at once.

JavaScript client: the pop defaults say what pop() does. POP_DEFAULTS and the README said a pop returns at once and that autoAck(true) commits it at delivery. Neither was ever true: a pop long-polls for timeoutMillis unless .wait(false), and autoAck() never reaches the broker, because it is the ack consume() sends after the handler. A pop commits at delivery only with the new commitOnDelivery(), below. POP_DEFAULTS.wait is now true, the README and the builder’s comments say so, and tests pin both. No behaviour changes.

JavaScript client: commitOnDelivery() commits a pop at delivery. The broker can move a consumer group’s cursor past the messages as it hands them out, with no lease and nothing to ack, but this client could not ask for it: autoAck(true) on a pop never reached the broker. queen.queue(q).group(g).commitOnDelivery().pop(), and popResult(), now send autoAck=true, the parameter every 2.x broker reads, and nothing else; the messages come back with an empty leaseId. This is at-most-once: a crash after the pop loses the messages. consume() always leases, so it throws before any request when the builder has commitOnDelivery(). autoAck() stays the ack consume() sends after the handler and still never reaches the broker. The broker refuses commitOnDelivery() together with conflation() (400), and pop() raises that 400. On ephemeral queues, queen.ephemeral.pop() takes commitOnDelivery: true; its autoAck option, which meant the same, still works and is deprecated, and passing both throws.

JavaScript client: admin.moveMessageToDLQ() and admin.clearQueue() are deprecated and throw. They sent POST /api/v1/messages/:partitionId/:transactionId/dlq and DELETE /api/v1/queues/:name/clear, routes the 2.x broker does not have, so every call failed with a 404 no_such_route and the message not found. Both now reject before any request and name the way that works. For a dead letter, ack the message with the dlq status, queen.ack(message, 'dlq', { group }), while its consumer group holds the lease. To skip what is queued, seek each consumer group to the end, queen.admin.seekConsumerGroup(group, queue, { toEnd: true }); a pop without a group reads as the group __QUEUE_MODE__.

PHP client 2.2.0 - 2026-10-06

Laravel: prefetch 'auto'. Each worker sizes its next pop from how long its jobs take, so a batch holds about 250 ms of work: short jobs get batches of up to 16, a job of a second or more gets one per pop. A job is timed from the moment the worker hands it to Laravel to its next pop, its ACK included, so an empty long poll or the worker’s sleep never counts as work. A queue starts at one job per pop and doubles only after two pops in a row came back full; a short pop sets the next one to what the queue had, an empty pop to one job, and slower jobs shrink it at once. A worker that serves several queues sizes each on its own. It runs in the worker, so it needs no supervisor and applies at once, and like any prefetch above 1 it needs lease_renewal. On the Linux server, 32 workers ran 2,003 jobs/s of 10 ms jobs with it against 1,548 at prefetch 1, and 2,867 with ack_async and pop_ahead, as many as a fixed prefetch of 4 with a third of its pops (benchmark-queen/2026-10-05-laravel-auto-prefetch). The Laravel guide now describes three profiles, safe, balanced and fast; balanced and fast long-poll (block_for 1) with the pool’s sleep at 0, since workers asleep when a burst began left it to the few awake ones and pushed the p99 at 500 jobs/s to 338 ms in one run of five. The dashboard shows 'auto', and its advice for short jobs suggests it.

PHP client 2.1.0, supervisor 0.7.0 - 2026-10-05

PHP client 2.1.0 pins supervisor 0.7.0: the Rust master reads lease_service from its configuration and handles the queue:restart exit below, so 0.6.0 cannot run with it.

Breaking, Laravel: config/queen.php reads 20 environment variables instead of 95. Only the values that differ between environments or deployments still read the environment: the broker URLs and token, the queue, consumer group and partitions, the supervisor’s read token, state directory and its remote status, prefork and coordination switches, the dashboard’s switch, path, domain and console URL, the alert mail, the metrics switch and token, and the supervisor binary’s install path and mirror. Every other setting is a plain value in the published file, with the default it had: set it there, or add your own env() call where a value must differ per environment. QUEEN_SUPERVISOR_LEASE_SERVICE, which the Rust master read from its environment, is the config key supervisor.lease_service (default on), so the dashboard now shows the master’s value instead of the web host’s. supervisor.remote_status.key is no longer required: it defaults to a slug of APP_NAME and APP_ENV. An application that publishes its own config/queen.php keeps every variable that file reads; the configuration reference maps each removed variable to its key.

Laravel prefork: php artisan queue:restart runs the deployed code. A forked worker stops at queue:restart like a spawned one, but its replacement was forked from the fork server, which still held the code it booted, so the workers kept the old code until the master restarted. A worker that stops for the restart signal now says so (Laravel 12 gives the reason), and both engines then start a new fork server: the workers forked from then on boot nothing and run the code on disk. The old server stays open until the last worker it forked exits. Changes to config/queen.php still need queen:supervisor terminate, as before.

Laravel prefork: a fork server that is not safe to fork is not used. fork() copies only the calling thread, so a booted application that runs another thread (a gRPC or Kafka extension, an APM agent) would give every worker a broken copy. On Linux the fork server now refuses to serve when another thread outlives a 5-second grace (libcurl’s resolver thread ends within it), names the threads, and the master spawns its workers instead. It warns about sockets the boot left open, which every forked worker would share, and it releases the database and Redis connections, log channels and mailers the boot opened before the first fork, not in each child only.

2.0.1 - 2026-10-05

Everything in 2.0.1-beta below, and:

Traces live on disk. A trace was a row in every node’s RAM, about one and a half times its stored size, until QUEEN_RAFT_TRACE_RETENTION_S (7 days by default) expired it, so a client that traced every message with a few KB of data grew every node by hundreds of MB an hour. Each node now keeps its traces in a second LMDB environment, <QUEEN_RAFT_DIR>/traces/, whose pages are page cache the kernel can drop: 100,000 traces of 5 KB add 56 MB to a node’s process memory instead of 810 MB. The trace routes, their fields, order and pagination are unchanged, and traces written before the upgrade stay readable until they expire. Expiry runs in bounded steps (512 traces each, a few ms at most) instead of one step for everything past the cutoff (139,000 traces took 1.24 s on 2026-10-05). New gauges: queen_raft_traces_map_bytes and queen_raft_traces_stored.

The upgrade is one-way. A cluster writes traces to disk once every member runs 2.0.1: the leader then raises the cluster version to 4, and from then on a 2.0.1-beta or older binary refuses to start on the data directory. On a cluster at version 4, a learner can be added only once its node is running and answering (as QUEEN_RAFT_JOIN already requires).

A trace with a very long transaction id or name no longer stops the cluster. Its store key went past LMDB’s 511-byte limit and every node stopped applying at that entry. The request is now refused with a 400 (name_too_long).

The dashboard shows the memory a node holds. A node’s memory meter showed its resident set, which counts the file pages the process maps: the store’s LMDB file is read whole at boot, so a node looked about 1 GB fuller than it was. The meter now shows the process’s anonymous memory, the part that can run a node out of memory, reported as anonBytes in /api/v1/raft/members; an older broker still shows its resident set.

2.0.1-beta - 2026-10-05

A new partition’s first message reaches every consumer group. On a queue read by two or more consumer groups, a group that took a new partition into the leader’s memory before the partition’s first message was written could miss that message until the partition’s next message or a change of leader. A second group loading the partition a moment later raised the shared tail and armed only its own cursor, so the append’s own wake-up found nothing left to do. A group’s load now arms every group already watching the partition whenever it finds the tail moved.

The leader keeps only the consumer state that is in use. The consumption engine kept every (consumer group, partition) pair it had served in the leader’s memory, about 1 KB each, until the partition or the group was deleted, and a partition is deleted only after PARTITION_CLEANUP_DAYS (30 by default). A workload that keeps opening partitions, one per conversation or entity, grew the leader without bound. A whole-queue group’s pair that holds nothing (no lease, nothing to deliver, no hold, every change durable) and stays so for QUEEN_CONSUME_IDLE_UNLOAD_S (default 600; 0 keeps every pair) now leaves memory. The next message on its partition loads it back from its cursor row, as a new leader does. New gauges: queen_consume_engine{kind="groups"|"parts"|"partitions"} and queen_consume_parts_total{kind="loaded"|"unloaded"}.

2.0.0 - 2026-10-03

The PostgreSQL storage class is removed. Queen 2.0 has one storage class, its own replicated log, and there is no database to run beside it. Each node keeps its whole state in one data directory, QUEEN_RAFT_DIR (default /var/lib/queen/raft): the queue logs that hold the payloads, an embedded ordered store for everything the broker looks up, and, on a cluster, the raft state. A single node answers a write once it is fsynced to its own disk. A cluster of three or five voters (QUEEN_RAFT_REPLICATOR=openraft, QUEEN_RAFT_NODE_ID, QUEEN_RAFT_PEERS, QUEEN_RAFT_TOKEN) has one leader order every write, answers it once a majority of the voters have it on disk, and keeps serving through the loss of a minority. QUEEN_STORAGE is gone with the choice it made. There is no in-place upgrade from 1.x: a 2.0 broker does not read a 1.x database and messages do not carry over, so a move is a cutover. Start 2.0 beside the old deployment, re-apply the queue configuration, move the producers, drain the old deployment and move the consumers.

The SQS facade is removed. It is not in the image or the binary any more, and QUEEN_SQS_* is not read.

The S3 sink runs inside the broker, on every node. QUEEN_S3_EMBEDDED=true starts the data-lake sink in the broker process, on a runtime of its own (QUEEN_S3_THREADS, default a quarter of the cores, 1 or 2); there is no queen-s3 binary, no child process, no QUEEN_S3_BIN and no /healthz or /metrics listener of its own, and QUEEN_S3_LISTEN and QUEEN_S3_LOG_FORMAT are named in a warning and ignored. Every node runs it with the same configuration, one sink per broker tenant (below). Per queue, a lease in the key/value store (s3:<sink>:<queue>:lease, TTL QUEEN_S3_LEASE_TTL_MS, default 30 s, refreshed every third of it) picks the node that writes the queue, and every window intent and commit carries that lease as a required conditional write, so a node that lost the queue cannot commit. A node claims a free queue after waiting one second for every queue it already runs or is claiming, plus a jitter under 200 ms, at most half the lease TTL, so the least-loaded node claims first and the queues spread over the nodes; a claim that fails for a passing reason (no leader yet) is retried within about a second rather than a TTL. Queues are then rebalanced: each sink counts the nodes running it through TTL’d presence rows, and a node holding more than ceil(queues / nodes) gives one queue back at a time (drained and released as at a SIGTERM) to a node below its share, so pods started one after another still end up sharing the queues, and nothing moves in a steady state. A node stopped with SIGTERM gives its queues back at once; one that dies loses them when its leases expire, and one that comes back under the same QUEEN_S3_INSTANCE (by default node-<id>@<host>) takes its own back at once. Each node reads its own applied copy of the log, followers included, and closes a window against its own safeTime; the lease refresh is a log entry, so safeTime keeps moving on a broker nothing writes to and the last window before a quiet spell still closes by age. The sink needs no QUEEN_URL and no token, and the proxy’s key/value carve-out for it is gone. Both record envelopes are the 1.5.0 formats; every key gains a tenant= level, the manifest a tenant field, the Parquet footer a queen.tenant pair, and the position checkpoint each partition’s id, so a 2.0 sink is pointed at a new prefix or bucket. A record’s ts is the stamp of the append that wrote it, and across the lake (partition, offset, ts) is unique, since a partition deleted and created again starts again at offset 0. A missing or bad value of one of the sink’s variables fails the broker’s boot, naming the variable; a bucket that does not answer delays the sink and nothing else. The window buffers are the broker’s memory now, so QUEEN_S3_MEMORY_MB defaults to 512 instead of 1024, and an out-of-memory takes the broker: size the container for both. On SIGTERM the sink stops reading at once, finishes the window it is committing and gives its leases back while the broker hands off and drains, and the broker waits for it up to QUEEN_S3_SHUTDOWN_GRACE_MS (30 s) counted from the signal. Its state is the s3 block of GET /status, and its queen_s3_* families are on /metrics/prometheus. queen_s3_lag_seconds is the node’s safeTime minus the stamp the queue’s lake is complete through (completeThrough in the status), so a queue read to its end and idle lags by about the guard and one discovery interval instead of growing; a node exports a queue’s gauges only while it runs the queue. The commit pointer’s tEnd is now an ISO-8601 timestamp: 1.5.0 wrote integer microseconds, which the retention hold could not read, so retentionSinkHold always sat at its cap; it now follows the sink. Running it is deploy/s3; what it writes is reference/s3.

Every broker tenant can have an S3 sink and a bucket of its own. The default tenant’s sink is configured by QUEEN_S3_* and turned on by QUEEN_S3_QUEUES: without it the default tenant has no sink, and another per-tenant QUEEN_S3_* variable set without it fails the boot. Every other tenant’s is set by the control plane: PUT /api/cp/clusters/:slug/s3 with the tenant’s endpoint, region, bucket, access key and queues, any other per-tenant setting, and secretKey, which is stored only sealed with the cell’s QUEEN_ENCRYPTION_KEY (the same on every node; a cell without one refuses the secret). GET answers it without the secret, DELETE removes it, and the tenant’s purge and delete remove it too. The routes are offered only where the node runs the sink. Every node reads those rows every 5 seconds and runs a tenant’s sink while its row is enabled and its cluster’s status, combined with its tenant’s, is active or push_blocked; a write that moves updatedAt, which every PUT with a secret does, rebuilds the sink once the old one has drained. The memory budget, threads, fetch concurrency, discovery interval, guard, lease TTL, multipart threshold, checkpoint cadence and instance name stay node-wide, in the environment, shared by every tenant’s sink, and a tenant’s settings cannot name them. Keys start with the tenant, <prefix>/tenant=<id>/queue=<name>/…, and the sidecars sit under <prefix>/_queen/tenant=<id>/queue=<name>/, so two tenants never share a key even in one bucket and prefix. A tenant’s leases and commit pointers are in its own key/value store, where its queues’ retention hold reads them. The s3 block of GET /status is {mode, phase, threads, controlPlane, sinks}, one entry per tenant with its source (env or cp), and a control-plane tenant’s queen_s3_* series carry a tenant label; the node renders each family once, and a removed tenant’s series go with it.

POST /api/v1/partitions/changed answers a sound safeTime. It is the greatest record stamp the answering node has applied: every record stamped at or below it is already readable on that node, so a time window that ends there is complete. 1.x derived it from PostgreSQL’s open transactions, with a fixed fallback floor (safeTimeDegraded, now always false), and the 2.0 betas answered the node’s wall clock, which a record planned before the answer and applied after it could fall below. It is per node, and it moves only when an entry applies, so an idle node answers the same value until the next write. Pages run in partition creation order with and without since, under an opaque cursor that means the same in both, and a complete pass reads each of the queue’s partitions once, where every page used to sort the whole queue; a cursor of the old shapes (n|…, t|…) is BAD_CURSOR. Every partition carries id, its uuid, which is new when a partition is deleted and created again under the same name. since is read to the microsecond, and one that is not a timestamp is a 400, as in 1.x, where the 2.0 betas read it as absent and listed everything.

The standalone proxy is removed. The proxy runs inside the broker process with QUEEN_PROXY_EMBEDDED=true, fronting PORT, or its own QUEEN_PROXY_PORT while PORT stays the internal broker port. Its tenants, clusters, users, API-key hashes, plans and usage live in the broker’s replicated key/value store under a system tenant, so the proxy database, PXDB_* and the proxy’s SQL migrations are gone, along with the queen-proxy binary and image. A first tenant can come from QUEEN_PROXY_BOOTSTRAP_* at boot, and the rest through the control plane under /api/cp/*, guarded by QUEEN_PROXY_CP_TOKEN.

The Kafka facade runs in-process. QUEEN_KAFKA_EMBEDDED=true starts it inside the broker, calling the broker’s router without a socket; there is no child process and no QUEEN_KAFKA_BIN. Committed offsets are Queen consumer-group positions by default (QUEEN_KAFKA_OFFSET_STORE, positions or kv).

Environment variables removed. Every PostgreSQL variable (PG_*, DB_POOL_SIZE, PG_USE_SSL, PG_SSL_REJECT_UNAUTHORIZED, PG_SSL_ROOT_CERT), the disk spool (FILE_BUFFER_*), the broker mesh (QUEEN_MESH_*, QUEEN_UDP_*, QUEEN_SYNC_*), the hot-list, fusion, ack-fusion, pop-fusion and admission knobs, QUEEN_APPLY_SCHEMA, RETENTION_PARALLELISM, the statistics refresh intervals, QUEEN_STORAGE, QUEEN_SQS_*, QUEEN_S3_BIN and PXDB_*. Some 1.x names stay because the 2.0 engine reads them: RETENTION_INTERVAL, RETENTION_BATCH_SIZE, PARTITION_CLEANUP_DAYS, QUEEN_PARTITION_CLEANUP_ENABLED, METRICS_FLUSH_MS, QUEEN_SWEEPER_BACKOFF_MIN_MS, QUEEN_SWEEPER_BACKOFF_MAX_MS, QUEEN_SWEEPER_TRANSIENT_BACKOFF_MS, QUEEN_SWEEPER_MAX_ATTEMPTS, QUEEN_STMT_TIMEOUT_MS, DEFAULT_TIMEOUT, POP_DEFAULT_TIMEOUT_MS and DEFAULT_SUBSCRIPTION_MODE.

Routes and answers. GET /api/v1/analytics/postgres-stats and the broker-to-broker /internal/api/* routes are gone, and /metrics no longer carries a database block. A push item answers queued, duplicate or error: with no spool there is no buffered. When a node cannot take a write the request fails with 503 and Retry-After, and a full disk answers 507. POST /api/v1/stats/refresh still answers 200 and does nothing, because counters are kept as entries apply.

Maintenance mode is removed. GET/POST /api/v1/system/maintenance, GET/POST /api/v1/system/maintenance/pop and GET /api/v1/status/buffers answer 404. No push or DLQ replay is refused with a maintenance 503, and no pop answers {"messages":[],"paused":true}. The embedded Rust API loses Broker::set_push_maintenance. The SDKs drop their maintenance calls (get/setMaintenanceMode and get/setPopMaintenanceMode in JS and PHP, the same four in Go, get/set_maintenance_mode and get/set_pop_maintenance_mode in Python, and Admin::maintenance, set_maintenance, pop_maintenance and set_pop_maintenance in Rust) along with their handling of a paused pop; queenctl maintenance and the dashboard’s maintenance switches are gone. The kv, timers and ephemeral kill switches and tenant quotas stay.

The embedded Rust API boots on a data directory. BrokerConfig::new().raft(dir), with raft_disk_pct(high, low) and stmt_timeout_ms; pg(), pg_use_ssl, pg_ssl_reject_unauthorized, pool_size, apply_schema, spool_dir, retention, stats_refresh, system_metrics and log_reports are removed. StartError has only Config, and Broker::shutdown() returns nothing.

One image, one binary. ghcr.io/queen-mq/queen carries the broker, with the proxy, the Kafka facade and the S3 sink linked in, the dashboard and queenctl; the queen-kafka, queen-sqs and queen-s3 binaries and the PostgreSQL client tools are gone. The dashboard is raft-only: the Postgres stats panel, the database pool and the disk-spool cards are removed, and the replicated log’s status takes their place.

Every client is 2.0.0. The JavaScript, Python and Rust packages (queen-mq on npm, PyPI and crates.io), the PHP client, the Go module, the C++ header and queenctl move to 2.0.0, with the Queen 2 broker. Go users change an import path: a major version above 1 is part of a Go module’s path, so the client is now github.com/smartpricing/queen/clients/client-go/v2 and the CLI installs with go install github.com/smartpricing/queen/clients/client-cli/v2/cmd/queenctl@latest. The tags keep their directory prefix (clients/client-go/v2.0.0, clients/client-cli/v2.0.0), and the client’s has to be pushed first, since the CLI requires it. queen-protocol stays at 1.3.0.

JavaScript, Python, Go, C++ and PHP clients, queenctl: every ACK of a transaction names its own lease. The transaction builders sent a message’s lease only in requiredLeases. A Queen 2 broker fences each ACK with the leaseId its operation carries, and lends it the one in requiredLeases only when the bundle names a single lease, so in a bundle that acked messages of two leases (two pops, two queues, a Laravel worker handing back two batches) every ACK went unfenced: it could complete a message that another consumer had taken after the lease expired, and commit the rest of the bundle with it. Each ACK operation now carries leaseId, and the broker refuses the whole transaction once any of those leases is no longer the caller’s. requiredLeases is still sent, and 1.x brokers already read the operation’s lease first. queenctl tx sent no lease at all: it dropped the bundle file’s requiredLeases. It now sends each ACK’s own leaseId, lends an ACK without one the lease requiredLeases names when it names exactly one, and exits 1 before sending anything when an ACK without a lease sits in a bundle that names several. The Rust client already named the lease on each ACK. PHP 1.9.0 shipped without this fix, its two-batch hand-back included.

Laravel dashboard: the Jobs page reads with a read-only token. It read the job metrics through POST /api/v1/kv, which takes read-write access, so with a read_bearer_token that may only read the page showed the metrics as unavailable on any broker that checks tokens. It now reads POST /api/v1/resources/kv/list, the read route the console’s KV browser uses (broker 1.6.0 and later), through the new Admin::listKv(). It falls back to getPrefix when the broker has no such route (404) and when the credential may not read but may use the KV surface (403), such as a proxy API key with consume and no read scope.

The partitions sunflower names its seeds. Each queue owns a wedge of the flower, and hovering a seed lights its queue and shows the queue, the partition, its pending and its lag. The per-partition figures come from a new read-only route, GET /api/v1/resources/partitions?queue=&limit=: the partitions holding the most pending (default 610), each with pending, processing and lagSeconds, the age of the oldest message its slowest reader has not consumed (null when caught up). The Overview asks for it only while it draws one seed per partition.

The broker in the top bar. For operators, every page’s top bar shows the cell’s CPU, memory, disk and Raft state, each figure its fullest node’s, over a hairline of how full it is; a click opens every node. Colour appears only past a line: 90% of the CPUs a node may use, 80% and 90% of its memory limit, and the node’s own disk write gate. GET /api/v1/raft/status and each member of GET /api/v1/raft/members gain host, the figures behind it.

The sidebar collapses to a rail of icons (the button at the top left of the top bar, or ⌘\), remembered per browser. Members and Users share one row: an operator switches between the acting cluster’s members and every account on the cell from the Members page.

PHP client and Laravel supervisor

Laravel: up to 1,024 stripes per queue. A stripe runs one job at a time, so the 64 stripes a queue could have capped its ordinary jobs at 64 busy workers. QUEEN_PARTITIONS now takes 1 to 1,024 (64 by default, unchanged), and a pop still asks for at most the 64 partitions the broker checks out per call: the broker serves the next ready ones in turn. On the Linux benchmark server, 128 workers drained 50,000 jobs at 9,257 jobs/s with 256 stripes against 4,722 with 64, 64 workers 20% faster, and the p95 at 500 jobs/s fell from 23.4 to 16.7 ms, with no measured cost at 50 jobs/s (benchmark-queen/2026-10-02-laravel-partition-stripes). Event-driven scaling watches up to 1,024 stripes per connection in one fetch; the regular poll finds jobs on the others.

Laravel: a crashed worker’s prefetched jobs keep their attempt. With prefetch above 1 and lease renewal, a worker that dies without a shutdown (SIGKILL, the kernel’s OOM killer, a PHP fatal error such as memory_limit) no longer leaves its unstarted jobs to lease expiry, which charged each one an attempt: with tries = 1 they failed without running, and a job that crashed its worker every time used up the attempts of the jobs prefetched with it. The worker now journals the transaction its shutdown would send, and rewrites one small record of it before each ACK or release. Whatever renews its lease sends that transaction once the worker has exited holding the lease, while the lease is still the worker’s: the Rust supervisor’s master on Linux, whose journal sits next to its lease socket, or the worker’s PHP lease helper, whose journal sits in a private temporary directory. Each unstarted job is completed and copied to its partition with the runs so far, the job that was running counts its run, and the jobs return at once instead of after retry_after. Every ACK in it names the lease, so the broker refuses the whole transaction once the lease is no longer the worker’s. Still charged one attempt: a lost node, a crash that takes the lease helper with the worker, and a batch popped ahead whose answer the worker had not read. The dashboard’s Configuration page, queen:supervise and the Rust supervisor (through queen:supervisor-config) now warn at start about a pool with tries 1 on a connection with prefetch above 1 or pop_ahead.

Laravel: a job handed back unstarted keeps its attempt. A worker that stopped with a prefetched tail (--memory, --max-jobs, a deploy, Laravel’s timeout handler), or that could not track a batch popped ahead, handed the jobs back with a retry ACK. The broker counts the next pop of such a position as a redelivery, so each hand-back charged an attempt to jobs that never ran, and with tries = 1 they failed with MaxAttemptsExceeded on their next delivery. The hand-back is now one transaction that completes each unstarted job and pushes a copy to its partition with the runs so far, as a release does. A crash is handled by the entry above.

Laravel: a forked child no longer kills its worker. With the PHP lease-renewal helper, a child forked by a job (Laravel’s fork concurrency driver, pcntl_fork()) stopped the parent’s helper when it exited, and the parent’s watchdog then SIGKILLed the parent mid-job; the child was also fenced as soon as a subprocess of its own ended.

PHP client: a detached request is written before the next job runs. With ack_async or pop_ahead, a request on a new connection (a worker’s first detached request, or one after the broker closed an idle keep-alive connection) was only started when it was sent: cURL wrote it when the request was settled, which with prefetch above 1 is after the next job. A hard kill during that job (shutdown grace exceeded, out of memory, node loss) lost the ACK, and the acknowledged job ran again when its lease expired. postDetached() and getDetached() now return once the whole request is written, the connection included; on the Guzzle transport (behind a proxy, or with QUEEN_SDK_HTTP_TRANSPORT=guzzle) this holds for a request with a body, the ACK. A request that cannot be written within the 5-second connect timeout throws at once, and the Laravel queue then acknowledges synchronously. A pop sent ahead, which carries no body, is waited for at most 250 ms, so a slow or dead backend does not hold up the next job.

PHP client: a process forked by a job exits. libcurl’s resolver threads do not survive fork(), and recent libcurl keeps them alive for a moment after each name resolution (2 seconds on the multi handle that carries detached requests). A child forked in that window, by Laravel’s fork concurrency driver or by pcntl_fork() in a job, waited for them forever when its exit freed the cURL handles it inherited, so the job hung until its timeout and left the child stuck. The cURL transport now sets CURLOPT_QUICK_EXIT and shares one DNS cache between its handles, so detached requests reuse the synchronous handle’s resolution and never start a resolver thread of their own.

Laravel: job metrics and tag records make one bounded attempt. They are written from every worker’s job events and on WorkerStopping, and used the queue’s ordinary client, whose retries held the worker on a slow or rate-limiting broker and could spend the shutdown grace before the prefetched tail was handed back. They now use one 2-second attempt.

Laravel dashboard: retry a failed job in one click. The failed-job page has a Retry now button. It runs queue:retry for that job, so the broker’s dead-letter entry and the failed_jobs row stay in step, and the page still shows the command for a terminal. Forgetting, flushing and pruning stay with Laravel’s commands. An exception during the retry is reported to the application’s log, and the page shows only its class: its message can quote the job’s payload. The id all is refused, since queue:retry all retries every failed job.

Laravel dashboard: what each queue holds now. The Workload page shows, for every supervised queue, the jobs waiting and running and how long the oldest unfinished job has waited, from one broker read per queue, cached for 5 seconds. QUEEN_DASHBOARD_CONSOLE_URL links each queue to the Queen console, which lists the messages themselves. An invalid value is reported to the application’s log and turns the links off; it never stops the application or its workers.

Laravel dashboard: the Configuration page is a tuning guide. It shows every resolved setting of the connection and the supervisor with its environment variable, each pool as the running supervisor published it, and advice from the supervisor’s state and the job metrics: job classes whose longest run exceeds shutdown_grace, one worker for several queues, short jobs on prefetch 1, prefork off, a lease helper per worker, polling instead of event-driven scaling, and pools that could run a job twice. A setting that can hold a credential shows only whether it is set. Job metrics now record each class’s longest run.

Laravel: a rejected broker URL no longer leaks its password. The supervisor configuration printed an invalid broker URL, credentials included, into logs and error pages; it now redacts them.

Supervisors (both engines). A worker that ran at least stable_after and exits non-zero (queue:work exits 12 at --memory, a job timeout kills the worker) is restarted at once: it no longer holds its pool at a single probe for stable_after. On stop, the workers get SIGTERM before the coordination leave and the remote status publish, which a slow broker could stretch past the platform’s stop deadline. The PHP engine’s event-driven watcher can no longer freeze the master loop on an answer that stalls mid-body, its fork server survives a failed pcntl_fork(), and preforked workers honour --quiet. The Rust engine trusts the platform’s CA store (and SSL_CERT_FILE), so a broker behind a private CA works, and its lease service fences a worker before it logs why, so a broken stderr pipe cannot skip the fence.

Supervisors: a job timeout is not a crash. Laravel SIGKILLs a worker whose job outlives its timeout, and both engines counted that exit as a crash: with job timeouts shorter than stable_after, a burst of them opened the restart circuit and left the pool at one probe worker for every job on its queue, the healthy ones included. The worker now leaves a marker in the state directory’s exits/ before it dies, and the master restarts it without backoff and without counting it. A worker that stops at --memory after it handled a job counts as a clean exit too; one that stops before any job still backs off, since its boot alone passes the limit. Every other non-zero exit counts as before, a SIGKILL without a marker included (the OOM killer, a lease fence, kill -9).

Navigation

Type to search…

↑↓ navigate↵ selectEsc close