---
title: "Supervisor status contract"
description: "Publish supervisor instances into Queen KV: group identity, instance slots, versioned documents, heartbeat expiry, and compatibility across implementations."
---

> Queen MQ documentation, for AI agents
> Complete self-contained summary of Queen MQ: https://queenmq.com/llms-brief.txt
> Fetch that first when the question is about the product rather than about this page.
> Index of all pages: https://queenmq.com/llms.txt

# Supervisor status contract

The broker dashboard discovers published supervisor status in the `queen-supervisor` KV namespace.
The publication envelope is independent of the publisher's language. The PHP and Rust engines
of the Laravel process supervisor use `queen.supervisor.status/v1`; all six SDKs and queenctl
use the [consumer document](#optional-sdk-consumer-supervision), `queen.consumer.status/v1`.
An application-owned scheduler can publish the consumer contract without a new dashboard
integration.

## Application overview

The dashboard shows one horizontal card per publication group, across all SDK languages
and the PHP/Rust process supervisors. Each card includes current worker capacity, pools
needing attention, heartbeat age and runtime-specific diagnostics: busy workers and cumulative
handler failures for SDK consumers, or draining workers and free process slots for process
supervisors. Expand the instances line to see every loaded instance, then
select an instance to open its queues and diagnostics. Search and engine filters select
whole groups: a matching healthy host never hides an unhealthy sibling. Display pagination
counts applications, while **Load more publications** extends the bounded KV scan.

A grey dot before the application's name means no issue was reported by the loaded
instances, as a healthy state is marked on every other page. An amber mark means attention
or incomplete status and a red one a critical issue, each with its label in the header; a
hollow dot means the instances are not running. The line under the card counts the instances
in each state: expand it to inspect all of them. A partial KV scan marks summaries as partial, even if the loaded instances
report no problems. The dashboard cannot infer missing replicas after their publications
have expired and disappeared.

Worker totals are available only when every loaded instance in the group has a fresh,
running observation and supplies the count. Shortfalls add per instance; excess workers
elsewhere do not cancel them. Queue names are deduplicated across replicas, dynamic consumer
scopes are identified separately, and shared backlog is never added. This aggregation
requires no publisher or broker protocol change.

Each card also reads the last hour of broker history for one selected, named queue through
`/api/v1/analytics/queue-ops`. **Traffic** shows incoming and delivered messages per minute;
**Backlog** shows sampled pending-message depth. Select another reported queue to change
the chart. The queue list deduplicates names and sums worker allocations across loaded
instances, while queue history and backlog are never multiplied by replica count.

These charts describe **all consumers of the selected queue on the current cluster**,
not the individual application's throughput. A supervisor may use another connection;
its publication does not establish that it consumes from this broker. No named queue or
unavailable analytics produces an explicit empty state. Missing samples remain gaps,
including incomplete traffic buckets, rather than implied zeroes. Current rates use the
most recent complete time bucket; ACK failures and backlog deltas use the available
samples in the hour. SDK handler failures are separate cumulative counters since each
instance started, not broker ACK failures or a historical error rate.

History reads are limited to the selected queues of displayed cards, with at most three
concurrent requests and one shared request per queue per refresh. Stable endpoint refusals
stop automatic analytics retries until an explicit refresh or a source/cluster switch.
The supervisor KV publication itself still contains only the latest observation, not
historical worker counts or per-instance throughput.

## Identity and keys

```text
namespace: queen-supervisor

pmsintool/<instance_id>/head
pmsintool/<instance_id>/chunk/0000
pmsintool/<instance_id>/chunk/0001

pmsintool-staging/<instance_id>/head
pmsintool-staging/<instance_id>/chunk/0000

coordination/v1/<scope>/<instance_id>
```

| Part | Contract |
| --- | --- |
| Namespace | `queen-supervisor` by convention. Custom namespaces remain supported through the dashboard's Source setting. |
| Group | The first segment, a stable application/deployment name. In Laravel this is `supervisor.remote_status.key`, whose default is a slug of `APP_NAME` and `APP_ENV`. |
| Group syntax | 1–255 ASCII letters, digits, dots, underscores or hyphens; the first character is a letter or digit. Case-sensitive; lowercase is recommended. No slashes or automatic normalization. The exact name `coordination` is reserved. |
| Instance | 16–128 lowercase hexadecimal characters. Generate an ID for each observed runtime lifetime: a master process, executed SDK consume loop or application-owned scheduler. Keep it for heartbeats; generate another when that runtime starts again. Hostname and PID are descriptive fields, not identity. |
| Slot | `<group>/<instance_id>`. Each publisher writes only its own slot. Different engines may share a group, but must not reuse an instance ID. |

The full identity is the acting tenant/cluster, namespace, group and instance ID. Give applications
and environments that share a store distinct groups, such as `pmsintool-production` and
`pmsintool-staging`. Replicas of the same deployment share the group. Groups with the same queue or
pool names remain separate. Scan `pmsintool/`, including the slash, to avoid matching
`pmsintool-staging/`; verify the decoded address belongs to that exact group as well.

Grouping organizes observations. It does not isolate queue data or define a consumer group.
Replica coordination uses the reserved `coordination/v1/` protocol: its scope depends on broker
endpoints, consumer group and queue set, independently of the publication group. Changing the
publication group therefore does not change autoscaling membership.

## One publication

The head value has this shape:

```json
{
  "format": "queen.supervisor.remote-status/v1",
  "write": "0123456789abcdef0123456789abcdef",
  "chunks": 1,
  "bytes": 1234
}
```

`bytes` is the actual length of the UTF-8 JSON document, at most 1,048,576 bytes. Split it into
chunks of at most 45,000 bytes, encode each as standard padded base64, and write its zero-based
index as a four-digit suffix. There are at most 24 chunks. Each chunk value contains `write`,
`index` and `data`; its `write` must match the head's new 32-character lowercase hexadecimal ID.

Write the head and every chunk in one KV batch with the same TTL. Publish again before the
heartbeat timeout. The current publishers default to the poll interval and a TTL of at least
300 seconds or twice the heartbeat timeout, whichever is larger. Old chunks may remain until
expiry; the current head and matching write IDs determine which chunks belong to the document.

Read all parts of a slot from one KV listing response. Validate format, write IDs, indexes, byte
length, UTF-8, schema and instance identity before displaying it. If a page ends partway through a
slot, read that entire slot on the next page instead of assembling different snapshots. Missing
parts, unknown versions and expired records cannot establish current worker health.

## Status document

The JSON inside the chunks declares `schema: "queen.supervisor.status/v1"`. Unknown extra fields
may be ignored. Breaking changes require a new schema or transport format version; readers must
not reinterpret an unsupported version as healthy.

| Fields | Meaning |
| --- | --- |
| `instance_id`, `hostname`, `pid` | Process identity and host information. The instance ID must equal the slot's ID. |
| `engine` | Implementation name, currently `php` or `rust`. This is an open string, not an enum or the workers' programming language. The broker dashboard accepts names up to 64 characters. |
| `updated_at_epoch` | Heartbeat time in integer Unix seconds. |
| `started_at_epoch`, `uptime_seconds` | Optional generation start in Unix seconds and non-negative integer uptime, measured with a monotonic clock at the heartbeat. Uptime is not extrapolated by the dashboard. |
| `engine_version`, `client_version` | Optional implementation and client package versions, up to 64 characters. PHP uses the installed Composer package version for both; Rust uses its build version and the PHP package version exported by Laravel. Missing versions remain unknown. |
| `state`, `paused`, `stopping` | Reported lifecycle state: `starting`, `running`, `paused`, `terminating` or `stopped`, with explicit booleans. |
| `ready`, `capacity_satisfied` | Master readiness and whether the desired capacity is present. |
| `configuration.heartbeat_timeout` | Positive number of seconds before the heartbeat is overdue. |
| `configuration.process_limit` | Positive total process budget. The v1 dashboard supports up to 4,096 processes. |
| `draining` | Total draining workers. |
| `pool_status` | At most 256 pool observations, identified by the pair `supervisor` and `queue`. |
| `process_budget` | `limit`, `used`, `available`, `active_worker_processes`, `draining_worker_processes`, `renewal_helpers_reserved`. |

Each pool reports `running`, `desired`, `draining`, `depth`, `depth_available`, `ready`, `capacity_satisfied` and
`healthy`. Restart evidence uses `restart_state` (`closed`, `backoff`, `open`, `probe`),
`restart_failures` and optional `restart_in_seconds`. Optional `pids` and `replicas` add host
and coordination detail. Budget evidence uses `process_cost_per_worker`, `reserved_processes`
and `renewal_helpers_reserved`; a worker with no helper costs one process and reserves zero
helpers.

Readiness and target capacity are separate observations: a ready pool may still be scaling toward
its target. Missing worker allocations are summed per pool; surplus workers in a different pool
do not cancel a shortfall. Draining workers and renewal helpers occupy the shared process budget.

Optional `configuration.supervisors` entries join to pool observations by `name` and an exact
member of `queues`. Ambiguous names are not joined. The dashboard displays `connection`,
`consumer_group`, `balance`, `strategy`, `min_processes`, `max_processes`, `timeout`, `retry_after`,
`lease_renewal`, `tries` and `memory`. Min/max apply to the named pool across its configured queues;
the queue row's running/desired counts describe that queue's allocation. Memory is a configured
limit in MB, not measured RSS. Zero attempts means unlimited attempts.

Counts are non-negative integers. The dashboard checks that the pool sums, draining total and
process budget agree before diagnosing a lack of process headroom. Missing or inconsistent
values remain unknown. Publishers must not invent zeroes or health flags for measurements they
do not have. The broker dashboard renders a selected set of fields, not arbitrary configuration;
publishers should still omit credentials from the document stored in KV.

## Freshness and compatibility

A recent heartbeat is a recent observation, not proof that a remote process is still alive. The
broker dashboard compares its read time with `updated_at_epoch` and the reported heartbeat timeout,
allowing five seconds of future clock skew. KV expiry takes precedence over a recent timestamp.
An expired record may remain visible until the broker sweeps it; a swept record disappears.
A failed read makes the previous result unconfirmed.

The optional runtime metadata does not change the v1 schema. Updated Rust engines request the
PHP client version with `QUEEN_SUPERVISOR_ENGINE_METADATA=1` during their internal
`queen:supervisor-config --for-engine` export. PHP adds the optional field only for that request,
preserving compatibility with older Rust engines that reject unknown configuration keys.

Queue throughput, acknowledgement failures and sampled backlog are separate broker observations,
not fields attributed to a supervisor instance. The pool drawer fetches them only on request, for
one queue on the acting cluster. They cover all consumers of that queue. The supervisor's named
connection may target another broker; a KV publication alone does not establish that mapping.

Legacy `<key>/head` and `<key>/chunk/NNNN` publications remain readable. Earlier keys containing
slashes are displayed under their complete legacy name, never merged into their first segment.
The reserved coordination root is excluded. Duplicate legacy and per-instance copies of the same
process are counted once within the same group.

New PHP and Rust publications reject invalid group names before writing. This is a best-effort
publication failure: it is logged without stopping local supervision. Rename an old nonconforming
`remote_status.key` on the publishers and any Laravel dashboard readers when upgrading. Existing
canonical names, including `pmsintool`, need no migration.

The namespace holds a latest status copy, not worker logs, durable history or remote commands.
Pause, continue and terminate remain host-local operations.

## Optional SDK consumer supervision

SDK consumer reporters use the same KV namespace, group syntax, instance slots,
atomic head/chunk publication and expiry rules above, with a separate document
schema: `queen.consumer.status/v1`. A consumer instance identifies one invocation
of a consume loop; multiple loops in one process have different instance IDs.
Consumers are internal tasks, not additional operating-system processes.

This lifetime applies to all six SDKs: an awaited JavaScript or Python consume
execution, Go `Consume(...).Execute(...)`, Rust `consume()` or `consume_batch()`,
C++ `consume()`, and PHP fluent `consume()->execute()`. `queenctl tail` uses one
instance per command invocation. Reusing a client, builder, queue or publication
group does not preserve identity or counters. A normal exit from a message or
idle limit ends that instance and publishes `stopped` with zero running workers.
Repeated short executions therefore leave multiple recent stopped observations
for the same hostname. SDK reporters do not aggregate them into process or pod
liveness. See [client APIs and timing differences](/start/clients/#consumer-supervision).

Reporting is explicitly enabled by the application and defaults to off. A disabled
reporter creates no timer, task, thread or KV traffic. The standard interval is 10
seconds, heartbeat timeout 30 seconds and record TTL 60 seconds. Cooperative PHP
uses a heartbeat timeout of `max(30, ceil(poll_timeout_ms / 1000) + 15)` seconds
(capped at 86400) and a TTL of at least twice that, so a normal long poll fits
inside its observation window. A publication is
best effort, serialized, bounded, and independent of message acknowledgement.
Publish a final `stopped` observation on orderly exit; after a crash, freshness and
TTL establish that current state is unknown. Publication failure must not stop
consumption or silently change its ACK, retry, cancellation or lease policy.

```json
{
  "schema": "queen.consumer.status/v1",
  "instance_id": "0123456789abcdef0123456789abcdef",
  "engine": "go",
  "execution_model": "goroutines",
  "hostname": "billing-1",
  "pid": 123,
  "state": "running",
  "updated_at_epoch": 1800000000,
  "started_at_epoch": 1799999900,
  "uptime_seconds": 100,
  "configuration": { "heartbeat_timeout": 30 },
  "pool_status": [{
    "name": "consumer",
    "queue": "orders",
    "namespace": null,
    "task": null,
    "consumer_group": "billing",
    "desired": 4,
    "running": 4,
    "busy": 1,
    "completed": 12,
    "failed": 2,
    "last_completed_at_epoch": 1799999995,
    "oldest_inflight_seconds": 20
  }]
}
```

`execution_model` is `async-tasks`, `goroutines`, `threads` or `cooperative`.
`state` is `starting`, `running`, `stopping` or `stopped`. Engine names identify the
publisher, not a process supervisor. `desired` is configured concurrency;
`running` counts loops that have not exited; `busy` counts active handler calls.
A publisher must observe exits rather than repeat the configured count forever.

`completed` and `failed` count successful and failed **handler calls** since this
consumer invocation began. A batch counts as one call. These are not ACK counters,
and an application callback that handles an error may return successfully.
`last_completed_at_epoch` records the last returned or failed handler, or `null`
when none has finished. `oldest_inflight_seconds` is elapsed monotonic time for
the oldest active handler, or `null` when none is active. Counters and durations
are non-negative integers. Timestamps use whole Unix seconds. The standard pool
name is `consumer`; queue may be `null` for a namespace/task selector. Queue-mode
consumers use `__QUEUE_MODE__` as their consumer group.

The dashboard keeps process budgets and restart diagnostics exclusive to process
supervisors. Consumer heartbeats do not assert broker readiness, ACK success or
handler progress. Event-loop starvation can stop asynchronous publication; a
cooperative publisher only reports when the consumer regains control, and may
become overdue during a long handler or network wait. Operators must compare
progress and in-flight age, and inspect application logs. SDK reporting does not
provide remote commands, process restart or autoscaling. No message payloads,
error text, credentials or arbitrary client configuration belong in these records.

### Reporting an application-owned scheduler

A scheduler may own persistent worker slots while using short SDK consume calls
to rotate between queues. To observe that runtime, implement a publisher of
`queen.consumer.status/v1` with the scheduler's lifetime. The SDK fluent
supervision options do not expose a shared reporter or an instance-ID override;
they describe individual consume executions. The same applies to a shell loop
that repeatedly starts `queenctl tail`: omit `--supervision-group` on its short
commands when publishing the outer scheduler's lifetime instead.

Generate one unique instance ID when the scheduler starts and keep it for every
heartbeat and its final status. Replicas share a publication group but use
different IDs. Generate a new ID after a restart. Leave per-call SDK supervision
disabled to avoid also publishing a separate instance for every turn.

Give each persistent worker slot a unique, stable pool `name`, such as
`worker-1`. A slot that processes one message at a time has `desired: 1`.
Report `running: 1` while its outer loop is alive, including idle polls and
queue changes, and `running: 0` after it exits. Report its currently assigned
`queue`, or `null` when none is assigned, and count only active handler calls in
`busy`. Counters accumulate for that slot across turns; they reset with the
new scheduler instance. Do not report the configured worker count as running
without observing whether those loops are still alive.

Use the atomic head/chunk publication above, a serialized heartbeat and bounded
best-effort writes. On shutdown, stop taking work, drain active handlers and
observe worker exits before the final `stopped` publication. A reporting failure
must leave scheduling, ACK/NACK, lease renewal and cancellation unchanged.
Retain limits used for fairness or a shared concurrency budget; replacing
short calls with unbounded per-queue consumers changes that scheduling policy.

Source: https://queenmq.com/reference/supervisor-status/index.mdx
