Every node serves Prometheus metrics in the text format on GET /metrics/prometheus. Each node
describes itself: what it served, how far its log and its apply have got, how long each stage of
its pipeline takes. So scrape every node of a cluster, and read a cluster as the set of its
nodes. Which of these to alert on, and at what thresholds, is on monitoring.
curl -s http://localhost:6632/metrics/prometheus | grep '^queen_raft_index'scrape_configs:
- job_name: queen
metrics_path: /metrics/prometheus
static_configs:
- targets: ["queen-0:6632", "queen-1:6632", "queen-2:6632"]On the broker port the route needs no authentication (it is in the default JWT_SKIP_PATHS). With
the embedded proxy on the public port (QUEEN_PROXY_EMBEDDED=true), /metrics and
/metrics/prometheus answer only a request carrying the control-plane token (x-queen-cp-token,
or Authorization: Bearer with the same value, set by QUEEN_PROXY_CP_TOKEN), and 404 to
everyone else, because they describe every tenant’s queues. The other way is to give the proxy its
own port with QUEEN_PROXY_PORT and scrape PORT from inside your network.
GET /metrics is a JSON summary for the dashboard: uptime, request and message totals, the
process’s resident memory, and a raft block with this node’s role, whether it knows a
leader, its term, applied and commit indexes and its lag. It is not Prometheus.
Reading the replicated log
queen_raft_index carries one series per kind: log (the last entry this node holds),
committed, applied and durable (the last entry its store checkpoint covers). The series
differ only by that label, so subtracting one from another needs ignoring(kind); without it
PromQL matches nothing and the expression stays empty.
# entries this node knows are committed but has not applied yet
queen_raft_index{kind="committed"} - ignoring(kind) queen_raft_index{kind="applied"}
# entries committed or proposed but not applied, as the node counts them itself
queen_raft_inflightFamilies
With QUEEN_RAFT_GROUPS above 1, each group has its own log and store, so the families that
describe them (queen_raft_index, queen_raft_inflight, queen_raft_log_storage,
queen_raft_proposals_total and the queen_raft_store_* families) carry a group label.
GET /metrics/prometheus exposes 107 families. Every one describes the node that answered: queen_process_* counts what this node did since it started, and the queen_raft_* families describe its replicated log, its store and its pipeline. Scrape every node. The pipeline timing, admission and forwarding families are skipped when QUEEN_RAFT_METRICS is 0, false, off or no; it is on by default. The per-queue families are written only for queues that have something to report. 27 families are declared but not written by this release; their Help says so.
Process (this broker instance)
| Family | Type | Help |
|---|---|---|
queen_event_loop_lag_avg_milliseconds |
gauge | Mean scheduler (event-loop) lag since start |
queen_malloc_bytes |
gauge | glibc heap: in use, free inside the arenas, arena total, mmapped chunks |
queen_parked_long_polls |
gauge | Currently parked long-poll pops on this process |
queen_process_ack_messages_total |
counter | Messages acked by this process |
queen_process_ack_requests_total |
counter | Ack API requests handled by this process |
queen_process_pop_messages_total |
counter | Messages popped by this process |
queen_process_pop_requests_total |
counter | Pop API requests handled by this process |
queen_process_push_messages_total |
counter | Messages pushed by this process |
queen_process_push_requests_total |
counter | Push API requests handled by this process |
queen_process_resident_memory_bytes |
gauge | Resident memory of this process |
queen_uptime_seconds |
gauge | Process uptime in seconds |
Replicated log and storage
| Family | Type | Help |
|---|---|---|
queen_raft_index |
gauge | Raft/RSM indexes |
queen_raft_inflight |
gauge | Entries committed or proposed but not applied |
queen_raft_log_storage |
gauge | Queue-log files and bytes |
queen_raft_proposals_total |
counter | Proposals accepted since open |
queen_raft_storage_full |
gauge | Disk/map admission gate |
queen_raft_store_commit_seconds |
summary | Store commit duration |
queen_raft_store_map_bytes |
gauge | LMDB map capacity and use |
queen_raft_store_operations_total |
counter | Embedded-store operations by kind |
queen_raft_store_ram_rows |
gauge | Rows per RAM keyspace: live, and keys dirty since the last checkpoint |
queen_raft_store_readers |
gauge | LMDB reader slots |
Admission
| Family | Type | Help |
|---|---|---|
queen_raft_admit_bytes |
gauge | Admission budget and bytes held by storage-growing commands in the pipeline |
queen_raft_admit_forwarded_total |
counter | Commands followers forwarded to this node that it admitted, and that it refused as overloaded |
queen_raft_admit_total |
counter | Commands that waited for admission room, and that were refused with 429 |
queen_raft_admit_waiting |
gauge | Commands waiting for admission room now |
Forwarding between nodes
| Family | Type | Help |
|---|---|---|
queen_raft_forward_total |
counter | Batched forwarding: commands this node sent to the leader, commands it served as the leader, streams it opened, and leaders it found without batched forwarding |
Pipeline timing
| Family | Type | Help |
|---|---|---|
queen_raft_apply_channel_depth |
summary | Apply channel depth at receive |
queen_raft_apply_entry_seconds |
summary | Per-entry apply duration (total) |
queen_raft_apply_other_seconds |
summary | Per-entry apply remainder (store puts + dispatch + derived) |
queen_raft_apply_receives_total |
counter | Apply-channel receives |
queen_raft_apply_segment_seconds |
summary | Per-entry segment write_all duration |
queen_raft_apply_stats |
counter | Apply thread counters (ApplyStats) |
queen_raft_arrival_to_proposed_seconds |
summary | Command arrival to its entry proposed |
queen_raft_checkpoint_rows |
summary | Rows per async checkpoint cut |
queen_raft_checkpoint_write_seconds |
summary | Async durable point: the checkpoint cut written, committed and synced (checkpoint thread) |
queen_raft_drain_commands |
summary | Drain size in commands |
queen_raft_drain_messages |
summary | Drain size in messages |
queen_raft_drain_to_propose_seconds |
summary | Cycle drain to its entry proposed (PERF-G leg 2) |
queen_raft_durable_dir_fsync_seconds |
summary | Durable point directory fsync duration |
queen_raft_durable_point_seconds |
summary | Durable point duration (total) |
queen_raft_durable_seg_fsync_seconds |
summary | Durable point segment-file fsync duration |
queen_raft_group_bytes |
summary | Log group size in bytes |
queen_raft_group_entries |
summary | Log group size in entries |
queen_raft_keep_overlay_total |
counter | Planning cycles by what the kept overlay did |
queen_raft_log_fsync_seconds |
summary | Log group fsync duration |
queen_raft_plan_cpu_seconds |
summary | Planner command loop, thread CPU time |
queen_raft_plan_deferred |
summary | Commands drained but deferred unplanned by the plan budget cut |
queen_raft_plan_seconds |
summary | Planner cycle duration |
queen_raft_plan_whole_cpu_seconds |
summary | Whole plan_cycle_blocking call, thread CPU |
queen_raft_plan_whole_wall_seconds |
summary | Whole plan_cycle_blocking call, wall |
queen_raft_planner_commands_total |
counter | Commands planned, by kind |
queen_raft_planner_plan_seconds |
summary | Per-command planning duration, by kind |
queen_raft_planner_slow_total |
counter | Slow-planned commands, by kind (O18) |
queen_raft_pop_h_total_seconds |
summary | PERF-J: whole raft pop handler (entry to response), pop-only |
queen_raft_pop_read_seconds |
summary | Pop payload read (segment render) latency |
queen_raft_propose_roundtrip_seconds |
summary | Entry proposed to its waiters answered (commit+apply+notify) |
queen_raft_proposed_to_committed_seconds |
summary | Entry proposed to its group fsynced |
queen_raft_push_h_await_seconds |
summary | PERF-J: push reply wait (propose+commit+apply+answer), push-only |
queen_raft_push_h_prep_seconds |
summary | PERF-J: push pre-submit (parse+pack), entry to first channel send |
queen_raft_push_h_submit_seconds |
summary | PERF-J: push channel enqueue (cmd_tx.send await), back-pressure |
queen_raft_push_h_total_seconds |
summary | PERF-J: whole raft push handler (entry to response), push-only |
queen_raft_qlog_early_codec_hits_total |
counter | Append blobs whose compression the facade started before planning |
queen_raft_qlog_early_codec_waiting |
gauge | Early compression jobs not yet taken by a proposal |
queen_raft_queue_wait_seconds |
summary | Command arrival to the cycle that drained it (PERF-G leg 1) |
queen_raft_slow_commands_total |
counter | Commands over QUEEN_RAFT_SLOW_COMMAND_MS |
queen_raft_writer_pickup_seconds |
summary | Entry proposed to the writer starting its group fsync (PERF-G) |
Per-queue rates and depth
| Family | Type | Help |
|---|---|---|
queen_dlq_depth_by_queue |
gauge | Dead letters per queue (exported by the raft leader only) |
queen_queue_conflated_per_minute |
gauge | Messages conflated away per queue in the last metrics bucket (METRICS_FLUSH_MS) |
queen_queue_pop_lag_milliseconds |
gauge | Per-queue pop lag (delivery time minus creation time) over the last metrics bucket (METRICS_FLUSH_MS), on this process |
Engine internals
| Family | Type | Help |
|---|---|---|
queen_batch_items_fired_total |
counter | Items flushed across fusion batches. Nothing in this release writes it, so it reads 0. |
queen_batch_rtt_milliseconds |
gauge | Fusion batch round-trip latency. Nothing in this release writes it, so it reads 0. |
queen_batches_fired_total |
counter | Fusion batches flushed. Nothing in this release writes it, so it reads 0. |
queen_fusion_items_per_batch |
gauge | Mean items per fusion batch. Nothing in this release writes it, so it reads 0. |
queen_pop_fill_wait_microseconds_total |
counter | Total time pops spent fattening an under-full batch. Nothing in this release writes it, so it reads 0. |
queen_pop_fill_wait_total |
counter | Pops that held an under-full batch back (minPopWaitTime). Nothing in this release writes it, so it reads 0. |
queen_pop_targeted_total |
counter | Hinted targeted single-partition pops issued. Nothing in this release writes it, so it reads 0. |
queen_pop_wildcard_total |
counter | Wildcard candidate-scan pops issued. Nothing in this release writes it, so it reads 0. |
Ephemeral queues
| Family | Type | Help |
|---|---|---|
queen_ephemeral_bytes |
gauge | Bytes held by this broker’s ephemeral rings |
queen_ephemeral_dropped_total |
counter | Ephemeral messages dropped, by cause |
queen_ephemeral_forwarded_total |
counter | Ephemeral requests relayed to the partition’s rendezvous owner |
queen_ephemeral_messages_total |
counter | Ephemeral messages by verb |
queen_ephemeral_queues |
gauge | Ephemeral queues on this broker (declared + live implicit) |
queen_ephemeral_wipes_total |
counter | Ephemeral rings dropped because their partition moved owner |
Key/value state, timers and the sweeper
| Family | Type | Help |
|---|---|---|
queen_kv_bytes_total |
counter | KV value bytes written / read |
queen_kv_expired_not_pruned |
gauge | Expired KV rows the sweeper has not pruned yet (capped). Nothing in this release writes it, so it reads 0. |
queen_kv_expired_not_pruned_capped |
gauge | 1 when the unpruned count hit its cap and is a floor. Nothing in this release writes it, so it reads 0. |
queen_kv_expiry_lag_seconds |
gauge | Age of the oldest expired, unpruned KV row. Nothing in this release writes it, so it reads 0. |
queen_kv_op_duration_milliseconds |
gauge | KV operation latency |
queen_kv_ops_total |
counter | KV operations by code path and outcome |
queen_kv_pool |
gauge | Dedicated KV connection pool. Nothing in this release writes it, so it reads 0. |
queen_kv_read_rejected_total |
counter | KV reads refused before reaching the database |
queen_kv_singleflight_coalesced_total |
counter | KV reads that shared an in-flight query. Nothing in this release writes it, so it reads 0. |
queen_sweeper_cycle_milliseconds |
gauge | Sweeper phase duration. Nothing in this release writes it, so it reads 0. |
queen_sweeper_phase_skipped_total |
counter | Phases shed under pressure (the degradation ladder, made visible). Nothing in this release writes it, so it reads 0. |
queen_sweeper_rows_total |
counter | Rows handled by each sweeper phase. Nothing in this release writes it, so it reads 0. |
queen_sweeper_skip_locked_total |
counter | Rows another broker was already holding. Nothing in this release writes it, so it reads 0. |
queen_sweeper_sleep_milliseconds |
gauge | Sleep the sweeper chose after the last cycle. Nothing in this release writes it, so it reads 0. |
queen_timers_dlq_total |
counter | Timers dead-lettered after exhausting attempts. Nothing in this release writes it, so it reads 0. |
queen_timers_due |
gauge | Timers due now, from the sweep probe (capped). Nothing in this release writes it, so it reads 0. |
queen_timers_due_capped |
gauge | 1 when the due count hit its cap and is a floor. Nothing in this release writes it, so it reads 0. |
queen_timers_fire_failures_total |
counter | Failed fire transactions by SQLSTATE class. Nothing in this release writes it, so it reads 0. |
queen_timers_fire_lag_seconds |
gauge | Delivery lateness of fired timers, per tenant. Nothing in this release writes it, so it has no samples. |
queen_timers_fire_lag_tenants_dropped_total |
counter | Fire-lag samples dropped because the tenant cap was reached. Nothing in this release writes it, so it reads 0. |
queen_timers_fired_total |
counter | Fired timer segments by outcome. Nothing in this release writes it, so it reads 0. |
queen_timers_oldest_late_seconds |
gauge | Lateness of the oldest due timer. Nothing in this release writes it, so it reads 0. |
queen_timers_poisoned_total |
counter | Batches replayed one segment per call after a permanent error. Nothing in this release writes it, so it reads 0. |
queen_timers_schedule_rejected_total |
counter | Timer schedules refused |
The per-queue families carry a queue label, and a tenant label for every tenant except the
default one; queen_queue_pop_lag_milliseconds adds stat (avg or max), measured over the last
metrics bucket (METRICS_FLUSH_MS, 60 s by default). queen_raft_forward_total is split by
kind: sent (commands this node sent to the leader), served (commands it served as the
leader), streams_opened, and legacy_leader (leaders it found that do not take batched
forwarding).
Some families come from earlier engines and nothing in 2.0 writes them, among them most of the
queen_timers_* and every queen_sweeper_* family: the table says so on each one, so do not build
an alert on them. A queen_raft_dbg family may also appear: those are temporary diagnostic
counters, and they may change or disappear in any release.