Unified observability plane

Global Command Centre

See metrics, events, logs and traces for the agent fleet in one place — then dig into evidence when something breaks. Read-only; no autonomous remediation.

Health

86

estate index

P1 open

2

escalated

Agents

30

14 high-risk

Read-only Agent OS
  • No shell execution
  • No cluster-admin
  • No secret reads
  • No database writes
  • No firewall changes
  • No autonomous remediation

Observability signals

Metrics · events · log pipelines · traces — unified fleet view
  • Nodes

    840

    infra

  • Clusters

    31

    infra

  • Agents

    30

    APM

  • Error events

    4

    events

  • P2

    2

    events

  • SLA risks

    3

    traces

  • Log drains

    4

    logs

  • Approvals

    5

    gates

Live telemetry

Fleet metrics & monitors

Streaming timeseries widgets and threshold monitors — Datadog-style dashboard layout, vendor-neutral simulated feed (1.5s ticks).

LIVEupdated now

CPU

41.7%

fleet avg

Latency p95

229ms

model gateway

Error rate

0.46%

production

Throughput

879rps

agent invoke

Agents busy

51%

active workers

CPU utilization

Agent hosts · last ~60s

Request latency

Gateway p95 · ms

Live event stream

Rolling ingest from collectors & agents

  • Collector heartbeat · eu-west-1

    23:23:26

  • Latency probe elevated · model gateway

    23:23:31

  • Collector heartbeat · eu-west-1

    23:23:34

  • Integrity check · evidence artefact

    23:23:30

Error rate

Failed invokes / total

Throughput

Requests per second

Active monitors

Threshold checks on live series — monitors-as-code pattern (no vendor lock-in)

  • Fleet CPU anomaly

    ok

    avg(last_5m):cpu.utilization{scope:agents}

    41.7% · thr > 85%

  • Gateway latency p95

    ok

    avg(last_5m):gateway.latency.p95

    229ms · thr > 400ms

  • Error rate spike

    ok

    sum(last_5m):errors.rate{env:production}

    0.46% · thr > 2.5%

  • Log pipeline lag

    ok

    avg(last_5m):pipeline.lag.p95

    1.7s · thr > 3s

Telemetry pipeline

Collectors forward structured container logs from agent sidecars — ops pattern analogous to Fluent Bit DaemonSets tailing /var/log/containers/*.log.

  • Log collectors (DaemonSet)

    4/4 nodes

    healthy
  • Structured JSON parse

    cri-o · containerd

    healthy
  • Export endpoint

    EU residency

    healthy
  • Pipeline lag p95

    1.4s · live

    healthy
Open evidence / log viewer

Application latency

Model gateway request performance — live p95 overlay on seeded providers

OpenAI

195 ms live · US / EU routing

healthy

Anthropic

225 ms live · US / EU routing

healthy

Google Gemini

255 ms live · US

degraded

Azure OpenAI

285 ms live · EU (Sweden Central)

healthy

AWS Bedrock

316 ms live · EU (Frankfurt)

healthy

Ollama (on-prem)

346 ms live · On-premise

healthy

vLLM Cluster

376 ms live · On-premise (sovereign)

healthy
Open Model Gateway

Infrastructure health heatmap

Composite health by customer and environment — see everything in one place

Customerproductionstagingdevdr
FS Core Banking Platform96918670
Nordic Payments Rail99998479
Card Issuing Services99938883
Grid Telemetry Fabric88837873
SCADA Edge Estate85807570
Clinical Data Platform93888378
Imaging AI Workloads99988277
National Registry Services96918670

Error & incident timeline

Historical event volume (seeded) — live series above for last-minute fleet health

Token spend and retry waste

USD per day across all tenants

Active incidents

Ordered by severity and SLA exposure

Recurring incident patterns

Signature clustering across 30 days

Registry egress reset during image pull

3 tenant(s) · last seen 2026-08-02

7x

Kafka consumer rebalance storm

2 tenant(s) · last seen 2026-07-30

5x

Replica lag during nightly ETL

1 tenant(s) · last seen 2026-08-01

4x

Edge collector telemetry drops

1 tenant(s) · last seen 2026-08-01

4x

Gateway config rollout latency regression

2 tenant(s) · last seen 2026-08-02

3x