What gets exported
Signals carry
service.name, version, supervision mode, and a stable one-way
hash of the install id so an operator can distinguish instances without
exporting account/profile identity or the raw install identifier.
mibyan.gateway.active_agents, mibyan.gateway.background_work, and
mibyan.gateway.background_delegations are complementary. active_agents
counts foreground message turns plus in-flight cron jobs plus API runs — the
work the gateway drains on shutdown. background_work counts detached work that
active_agents never includes: backgrounded delegate_task subagents,
terminal(background=true) processes, and kanban workers; it is
task-granular — a fan-out batch of N subagents counts as N — so it reflects
real concurrent subagent load. background_delegations counts only async
delegation units (each delegate_task dispatch is one, a fan-out batch is
one), matching the async pool’s capacity accounting; alert it against
delegation.max_concurrent_children to see slot pressure. Sum active_agents
and background_work for total live work per instance; use
background_delegations for pool-saturation.
Enabling
otlp extra and installs on first use
when policy permits. To request it explicitly from a prepared checkout, run
python -c "import pm; pm.sync_venv(['otlp'], explicit=True)". When the SDK is missing or the endpoint is down,
the gateway runs unaffected: metric collection and ordinary event export stay
off the hot path, while terminal cron events make one bounded fail-open flush
attempt of up to one second so the final state is less likely to be lost.
Works identically under systemd/launchd/s6 supervision, containers, tmux, or
a plain mibyan gateway run: the exporter lives in the gateway process, so
no sidecar, agent, or collector is required on the host.
Collecting into DataDog
Run a customer-owned OpenTelemetry Collector and forward:monitoring.export.otlp.endpoint at the collector. Alerts belong on
mibyan.gateway.up, mibyan.platform.up, and mibyan.platform.degraded.
Generic fleet queries and alerts
The exact syntax depends on the customer’s observability backend. The examples below use PromQL-style expressions and intentionally avoid vendor-specific routing, destinations, or customer inventory. Group fleet views by the opaqueservice.instance.id resource attribute. A
process that has died cannot emit its own zero, so every deployment needs both
explicit-state and missing-series detection.
mibyan.cron_execution spans.
Alert or derive events from bounded attributes such as:
- one row per
service.instance.idwith gateway and configured local-platform state; - scheduler heartbeat, last-success age, running count, overdue count, and catch-up increase;
- a cron lifecycle feed keyed only by opaque
mibyan.job_key; - separate alerts for box absence, local bridge down, scheduler stale, cron failed/unknown, delivery failure, and overdue/catch-up activity.
Release-validation scenarios
Before accepting a deployment, force and verify all five cases through the real collector and backend:- Cron success: observe
claimed -> running -> completed, duration, and a truthful delivery outcome. - Cron failure: observe
failedplus a bounded error class, with no raw exception or content in the decoded OTLP payload. - Cron interruption: stop the owning gateway during execution, restart it,
and observe recovery to
unknown. - Locally owned bridge outage: break one native connector, observe its bounded down/retrying/fatal state and recovery, and verify unaffected boxes remain healthy.
- Killed gateway: terminate one canary, verify missing-series detection, restart it, and confirm the same opaque instance identity returns.
Local smoke test (no Docker)
Maintaining and extending this plane
This plane is a fixed, enumerated, content-free vocabulary by design. Adding a signal is not just “emit a new metric” — every new name and attribute must be declared in each layer that enforces the bounded vocabulary, or it is silently dropped downstream. Follow the checklist for the change you are making. The golden rule: a new signal that is emitted but not declared in every layer looks like a code bug but is a vocabulary-registration bug — nothing errors, the signal just never arrives.Content-free invariant (applies to every change)
Before adding anything, confirm it cannot carry content. Numbers, booleans, ages, durations, monotonic counts, and one-way hashes are safe. Never add an attribute that can hold a job name, prompt, output, schedule, destination, raw exception text, file path, profile name, account id, or free-form string. When you must key a record to a job/entity, hash it (sha256(...)[:24], see
_job_key in agent/monitoring/cron_health.py) — never emit the raw id. All
string attributes that could touch user input must pass through
redaction.redact_for_export and be truncated (see _span_attrs in
agent/monitoring/otlp_exporter.py).
Adding a new gauge/metric
- Emit it in the snapshot builder (
agent/monitoring/gateway_health.pybuild_gateway_health_snapshot,cron_health.pybuild_cron_health_snapshot, or a sibling reader wired into_read_runtime_snapshotingateway_health_export.py). Best-effort: never let a reader raise into the collection loop — wrap it and log a content-free WARNING with the exception TYPE name only (the pattern the cron and background-work readers use), so a future regression is visible instead of silently dropping the signal. - Register the dotted metric name in the observable-gauge
metric_nameslist ingateway_health_export.py::_start_metric_provider. A gauge that is emitted in the snapshot but not registered here is never observed. - Add the export-table row and an alert example in this file.
- If the deployment fronts the exporter with an OpenTelemetry Collector that
uses a metric-name allowlist (a
filter/...processor withname != "..."guards), add the new name there too — otherwise the collector drops it before the backend. This is not repo code, but it is the single most common reason a correctly-emitted new metric never appears; call it out in the PR so the deploying operator updates their collector config.
Adding a new subsystem (a new family of signals)
Mirror the cron pattern (cron_health.py + its wiring): put the read/projection
logic in its own module, expose one build_<subsystem>_health_snapshot() that
returns bounded GatewayMetrics (and events if any), and extend it into
_read_runtime_snapshot with the same best-effort try/except-WARNING guard.
Then do the “adding a metric” checklist for each new name, and the “adding an
attribute” checklist for each new event attribute. Add a release-validation
scenario below for the subsystem’s failure mode.
Extending the error-class / status / source / state vocabularies
These are the closed enums that keep the plane bounded. Extend the SET, then the classifier, never one without the other:- Cron (
agent/monitoring/cron_health.py):_KNOWN_STATUSES,_KNOWN_SOURCES,_KNOWN_DELIVERY_OUTCOMES, and theclassify_cron_errorkeyword buckets. Anything not in the set is coerced tounknownon the way out, so a new value that is not added to the set is invisible. - Gateway/platform (
agent/monitoring/gateway_health.py):_KNOWN_GATEWAY_STATES,_KNOWN_PLATFORM_STATES, andclassify_gateway_error.
mibyan.error_class = ... list in this file’s alert section and the enum’s unit
test so the contract is asserted, not frozen as a count.
Adding a content-free attribute to an existing event/span
Add the key to the emitter’s per-kindkeep_by_kind allowlist in
agent/monitoring/otlp_exporter.py::_span_attrs (unlisted keys are dropped), run
it through redaction if it is ever string-shaped, and — as with metrics — if the
deployment’s collector has a span-attribute keep_keys(...) allowlist, add the
attribute there too or it is stripped in transit.
Verify the whole chain, not just emission
Emitting is necessary but not sufficient. Confirm the signal survives all the way to the backend, because the enums, themetric_names registration, the
emitter attribute allowlist, and any collector allowlist each drop unlisted
values with no error:
Boundaries and roadmap
Themibyan monitoring CLI intentionally exposes status only. This first
release covers only Mibyan-owned service-health and operational-diagnostic
signals, including Mibyan-owned Relay transport health. Team Gateway’s
authoritative shared connector/platform state is explicitly out of scope, as
are product analytics, audit/quality reporting, and detailed execution traces.
Shared client usage metrics and enterprise trace telemetry are being designed on
the NeMo Relay integration with their own consent, policy, and export
boundaries; this monitoring plane stays narrow so an operator can enable it
without touching any content-bearing signal. The telemetry surface may be
reorganized as that lands.
