Skip to content

Supervisor status contract

Publish supervisor instances into Queen KV: group identity, instance slots, versioned documents, heartbeat expiry, and compatibility across implementations.

Updated View as Markdown

The broker dashboard discovers published supervisor status in the queen-supervisor KV namespace. The publication envelope is independent of the publisher’s language. The PHP and Rust engines of the Laravel process supervisor use queen.supervisor.status/v1; all six SDKs and queenctl use the consumer document, queen.consumer.status/v1. An application-owned scheduler can publish the consumer contract without a new dashboard integration.

Application overview

The dashboard shows one horizontal card per publication group, across all SDK languages and the PHP/Rust process supervisors. Each card includes current worker capacity, pools needing attention, heartbeat age and runtime-specific diagnostics: busy workers and cumulative handler failures for SDK consumers, or draining workers and free process slots for process supervisors. Expand the instances line to see every loaded instance, then select an instance to open its queues and diagnostics. Search and engine filters select whole groups: a matching healthy host never hides an unhealthy sibling. Display pagination counts applications, while Load more publications extends the bounded KV scan.

A grey dot before the application’s name means no issue was reported by the loaded instances, as a healthy state is marked on every other page. An amber mark means attention or incomplete status and a red one a critical issue, each with its label in the header; a hollow dot means the instances are not running. The line under the card counts the instances in each state: expand it to inspect all of them. A partial KV scan marks summaries as partial, even if the loaded instances report no problems. The dashboard cannot infer missing replicas after their publications have expired and disappeared.

Worker totals are available only when every loaded instance in the group has a fresh, running observation and supplies the count. Shortfalls add per instance; excess workers elsewhere do not cancel them. Queue names are deduplicated across replicas, dynamic consumer scopes are identified separately, and shared backlog is never added. This aggregation requires no publisher or broker protocol change.

Each card also reads the last hour of broker history for one selected, named queue through /api/v1/analytics/queue-ops. Traffic shows incoming and delivered messages per minute; Backlog shows sampled pending-message depth. Select another reported queue to change the chart. The queue list deduplicates names and sums worker allocations across loaded instances, while queue history and backlog are never multiplied by replica count.

These charts describe all consumers of the selected queue on the current cluster, not the individual application’s throughput. A supervisor may use another connection; its publication does not establish that it consumes from this broker. No named queue or unavailable analytics produces an explicit empty state. Missing samples remain gaps, including incomplete traffic buckets, rather than implied zeroes. Current rates use the most recent complete time bucket; ACK failures and backlog deltas use the available samples in the hour. SDK handler failures are separate cumulative counters since each instance started, not broker ACK failures or a historical error rate.

History reads are limited to the selected queues of displayed cards, with at most three concurrent requests and one shared request per queue per refresh. Stable endpoint refusals stop automatic analytics retries until an explicit refresh or a source/cluster switch. The supervisor KV publication itself still contains only the latest observation, not historical worker counts or per-instance throughput.

Identity and keys

namespace: queen-supervisor

pmsintool/<instance_id>/head
pmsintool/<instance_id>/chunk/0000
pmsintool/<instance_id>/chunk/0001

pmsintool-staging/<instance_id>/head
pmsintool-staging/<instance_id>/chunk/0000

coordination/v1/<scope>/<instance_id>
Part Contract
Namespace queen-supervisor by convention. Custom namespaces remain supported through the dashboard’s Source setting.
Group The first segment, a stable application/deployment name. In Laravel this is supervisor.remote_status.key, whose default is a slug of APP_NAME and APP_ENV.
Group syntax 1–255 ASCII letters, digits, dots, underscores or hyphens; the first character is a letter or digit. Case-sensitive; lowercase is recommended. No slashes or automatic normalization. The exact name coordination is reserved.
Instance 16–128 lowercase hexadecimal characters. Generate an ID for each observed runtime lifetime: a master process, executed SDK consume loop or application-owned scheduler. Keep it for heartbeats; generate another when that runtime starts again. Hostname and PID are descriptive fields, not identity.
Slot <group>/<instance_id>. Each publisher writes only its own slot. Different engines may share a group, but must not reuse an instance ID.

The full identity is the acting tenant/cluster, namespace, group and instance ID. Give applications and environments that share a store distinct groups, such as pmsintool-production and pmsintool-staging. Replicas of the same deployment share the group. Groups with the same queue or pool names remain separate. Scan pmsintool/, including the slash, to avoid matching pmsintool-staging/; verify the decoded address belongs to that exact group as well.

Grouping organizes observations. It does not isolate queue data or define a consumer group. Replica coordination uses the reserved coordination/v1/ protocol: its scope depends on broker endpoints, consumer group and queue set, independently of the publication group. Changing the publication group therefore does not change autoscaling membership.

One publication

The head value has this shape:

{
  "format": "queen.supervisor.remote-status/v1",
  "write": "0123456789abcdef0123456789abcdef",
  "chunks": 1,
  "bytes": 1234
}

bytes is the actual length of the UTF-8 JSON document, at most 1,048,576 bytes. Split it into chunks of at most 45,000 bytes, encode each as standard padded base64, and write its zero-based index as a four-digit suffix. There are at most 24 chunks. Each chunk value contains write, index and data; its write must match the head’s new 32-character lowercase hexadecimal ID.

Write the head and every chunk in one KV batch with the same TTL. Publish again before the heartbeat timeout. The current publishers default to the poll interval and a TTL of at least 300 seconds or twice the heartbeat timeout, whichever is larger. Old chunks may remain until expiry; the current head and matching write IDs determine which chunks belong to the document.

Read all parts of a slot from one KV listing response. Validate format, write IDs, indexes, byte length, UTF-8, schema and instance identity before displaying it. If a page ends partway through a slot, read that entire slot on the next page instead of assembling different snapshots. Missing parts, unknown versions and expired records cannot establish current worker health.

Status document

The JSON inside the chunks declares schema: "queen.supervisor.status/v1". Unknown extra fields may be ignored. Breaking changes require a new schema or transport format version; readers must not reinterpret an unsupported version as healthy.

Fields Meaning
instance_id, hostname, pid Process identity and host information. The instance ID must equal the slot’s ID.
engine Implementation name, currently php or rust. This is an open string, not an enum or the workers’ programming language. The broker dashboard accepts names up to 64 characters.
updated_at_epoch Heartbeat time in integer Unix seconds.
started_at_epoch, uptime_seconds Optional generation start in Unix seconds and non-negative integer uptime, measured with a monotonic clock at the heartbeat. Uptime is not extrapolated by the dashboard.
engine_version, client_version Optional implementation and client package versions, up to 64 characters. PHP uses the installed Composer package version for both; Rust uses its build version and the PHP package version exported by Laravel. Missing versions remain unknown.
state, paused, stopping Reported lifecycle state: starting, running, paused, terminating or stopped, with explicit booleans.
ready, capacity_satisfied Master readiness and whether the desired capacity is present.
configuration.heartbeat_timeout Positive number of seconds before the heartbeat is overdue.
configuration.process_limit Positive total process budget. The v1 dashboard supports up to 4,096 processes.
draining Total draining workers.
pool_status At most 256 pool observations, identified by the pair supervisor and queue.
process_budget limit, used, available, active_worker_processes, draining_worker_processes, renewal_helpers_reserved.

Each pool reports running, desired, draining, depth, depth_available, ready, capacity_satisfied and healthy. Restart evidence uses restart_state (closed, backoff, open, probe), restart_failures and optional restart_in_seconds. Optional pids and replicas add host and coordination detail. Budget evidence uses process_cost_per_worker, reserved_processes and renewal_helpers_reserved; a worker with no helper costs one process and reserves zero helpers.

Readiness and target capacity are separate observations: a ready pool may still be scaling toward its target. Missing worker allocations are summed per pool; surplus workers in a different pool do not cancel a shortfall. Draining workers and renewal helpers occupy the shared process budget.

Optional configuration.supervisors entries join to pool observations by name and an exact member of queues. Ambiguous names are not joined. The dashboard displays connection, consumer_group, balance, strategy, min_processes, max_processes, timeout, retry_after, lease_renewal, tries and memory. Min/max apply to the named pool across its configured queues; the queue row’s running/desired counts describe that queue’s allocation. Memory is a configured limit in MB, not measured RSS. Zero attempts means unlimited attempts.

Counts are non-negative integers. The dashboard checks that the pool sums, draining total and process budget agree before diagnosing a lack of process headroom. Missing or inconsistent values remain unknown. Publishers must not invent zeroes or health flags for measurements they do not have. The broker dashboard renders a selected set of fields, not arbitrary configuration; publishers should still omit credentials from the document stored in KV.

Freshness and compatibility

A recent heartbeat is a recent observation, not proof that a remote process is still alive. The broker dashboard compares its read time with updated_at_epoch and the reported heartbeat timeout, allowing five seconds of future clock skew. KV expiry takes precedence over a recent timestamp. An expired record may remain visible until the broker sweeps it; a swept record disappears. A failed read makes the previous result unconfirmed.

The optional runtime metadata does not change the v1 schema. Updated Rust engines request the PHP client version with QUEEN_SUPERVISOR_ENGINE_METADATA=1 during their internal queen:supervisor-config --for-engine export. PHP adds the optional field only for that request, preserving compatibility with older Rust engines that reject unknown configuration keys.

Queue throughput, acknowledgement failures and sampled backlog are separate broker observations, not fields attributed to a supervisor instance. The pool drawer fetches them only on request, for one queue on the acting cluster. They cover all consumers of that queue. The supervisor’s named connection may target another broker; a KV publication alone does not establish that mapping.

Legacy <key>/head and <key>/chunk/NNNN publications remain readable. Earlier keys containing slashes are displayed under their complete legacy name, never merged into their first segment. The reserved coordination root is excluded. Duplicate legacy and per-instance copies of the same process are counted once within the same group.

New PHP and Rust publications reject invalid group names before writing. This is a best-effort publication failure: it is logged without stopping local supervision. Rename an old nonconforming remote_status.key on the publishers and any Laravel dashboard readers when upgrading. Existing canonical names, including pmsintool, need no migration.

The namespace holds a latest status copy, not worker logs, durable history or remote commands. Pause, continue and terminate remain host-local operations.

Optional SDK consumer supervision

SDK consumer reporters use the same KV namespace, group syntax, instance slots, atomic head/chunk publication and expiry rules above, with a separate document schema: queen.consumer.status/v1. A consumer instance identifies one invocation of a consume loop; multiple loops in one process have different instance IDs. Consumers are internal tasks, not additional operating-system processes.

This lifetime applies to all six SDKs: an awaited JavaScript or Python consume execution, Go Consume(...).Execute(...), Rust consume() or consume_batch(), C++ consume(), and PHP fluent consume()->execute(). queenctl tail uses one instance per command invocation. Reusing a client, builder, queue or publication group does not preserve identity or counters. A normal exit from a message or idle limit ends that instance and publishes stopped with zero running workers. Repeated short executions therefore leave multiple recent stopped observations for the same hostname. SDK reporters do not aggregate them into process or pod liveness. See client APIs and timing differences.

Reporting is explicitly enabled by the application and defaults to off. A disabled reporter creates no timer, task, thread or KV traffic. The standard interval is 10 seconds, heartbeat timeout 30 seconds and record TTL 60 seconds. Cooperative PHP uses a heartbeat timeout of max(30, ceil(poll_timeout_ms / 1000) + 15) seconds (capped at 86400) and a TTL of at least twice that, so a normal long poll fits inside its observation window. A publication is best effort, serialized, bounded, and independent of message acknowledgement. Publish a final stopped observation on orderly exit; after a crash, freshness and TTL establish that current state is unknown. Publication failure must not stop consumption or silently change its ACK, retry, cancellation or lease policy.

{
  "schema": "queen.consumer.status/v1",
  "instance_id": "0123456789abcdef0123456789abcdef",
  "engine": "go",
  "execution_model": "goroutines",
  "hostname": "billing-1",
  "pid": 123,
  "state": "running",
  "updated_at_epoch": 1800000000,
  "started_at_epoch": 1799999900,
  "uptime_seconds": 100,
  "configuration": { "heartbeat_timeout": 30 },
  "pool_status": [{
    "name": "consumer",
    "queue": "orders",
    "namespace": null,
    "task": null,
    "consumer_group": "billing",
    "desired": 4,
    "running": 4,
    "busy": 1,
    "completed": 12,
    "failed": 2,
    "last_completed_at_epoch": 1799999995,
    "oldest_inflight_seconds": 20
  }]
}

execution_model is async-tasks, goroutines, threads or cooperative. state is starting, running, stopping or stopped. Engine names identify the publisher, not a process supervisor. desired is configured concurrency; running counts loops that have not exited; busy counts active handler calls. A publisher must observe exits rather than repeat the configured count forever.

completed and failed count successful and failed handler calls since this consumer invocation began. A batch counts as one call. These are not ACK counters, and an application callback that handles an error may return successfully. last_completed_at_epoch records the last returned or failed handler, or null when none has finished. oldest_inflight_seconds is elapsed monotonic time for the oldest active handler, or null when none is active. Counters and durations are non-negative integers. Timestamps use whole Unix seconds. The standard pool name is consumer; queue may be null for a namespace/task selector. Queue-mode consumers use __QUEUE_MODE__ as their consumer group.

The dashboard keeps process budgets and restart diagnostics exclusive to process supervisors. Consumer heartbeats do not assert broker readiness, ACK success or handler progress. Event-loop starvation can stop asynchronous publication; a cooperative publisher only reports when the consumer regains control, and may become overdue during a long handler or network wait. Operators must compare progress and in-flight age, and inspect application logs. SDK reporting does not provide remote commands, process restart or autoscaling. No message payloads, error text, credentials or arbitrary client configuration belong in these records.

Reporting an application-owned scheduler

A scheduler may own persistent worker slots while using short SDK consume calls to rotate between queues. To observe that runtime, implement a publisher of queen.consumer.status/v1 with the scheduler’s lifetime. The SDK fluent supervision options do not expose a shared reporter or an instance-ID override; they describe individual consume executions. The same applies to a shell loop that repeatedly starts queenctl tail: omit --supervision-group on its short commands when publishing the outer scheduler’s lifetime instead.

Generate one unique instance ID when the scheduler starts and keep it for every heartbeat and its final status. Replicas share a publication group but use different IDs. Generate a new ID after a restart. Leave per-call SDK supervision disabled to avoid also publishing a separate instance for every turn.

Give each persistent worker slot a unique, stable pool name, such as worker-1. A slot that processes one message at a time has desired: 1. Report running: 1 while its outer loop is alive, including idle polls and queue changes, and running: 0 after it exits. Report its currently assigned queue, or null when none is assigned, and count only active handler calls in busy. Counters accumulate for that slot across turns; they reset with the new scheduler instance. Do not report the configured worker count as running without observing whether those loops are still alive.

Use the atomic head/chunk publication above, a serialized heartbeat and bounded best-effort writes. On shutdown, stop taking work, drain active handlers and observe worker exits before the final stopped publication. A reporting failure must leave scheduling, ACK/NACK, lease renewal and cancellation unchanged. Retain limits used for fairness or a shared concurrency budget; replacing short calls with unbounded per-queue consumers changes that scheduling policy.

Source of truth
Navigation

Type to search…

↑↓ navigate↵ selectEsc close