Skip to content

4 · Observe

Your system is live. This stage is about the views that answer the daily questions — what is it doing, what does it cost, what got blocked — and all of them share one property: they filter to your AI System. You are never reading a firehose of everyone’s traffic and guessing which lines are yours.

Mission Control is the operations home of the Admin UI, with three entries in the sidebar:

  • Overview — the estate as Resource Group cards. Drill from your team’s card to your system’s, or search for it: spend, runs and cost per completed task for the selected window, and — pinned to the standard 30-day assurance window regardless of the selector — its health, its assurance verdict and its Big-T workload class. A red badge on the card counts its open findings.
  • Analytics — the tabbed metrics view. Open it from your card’s Metrics action, or pick your system in the scope selector at the top, and every tab re-scopes. Choose a timespan (24h → all time) and, if you import external telemetry, whether you’re looking at governed traffic, observed traffic, or both.
  • Inbox — the findings that need a human decision, covered under Alerts below.

The Analytics tabs you’ll live in:

Tab The question it answers
Traffic Is it up, is it erroring, what’s the request volume — and which users and tools are driving it?
Cost Total spend, attributed by provider, model, group and user — with a residual line for agent/unattributed traffic so per-user spend reconciles to the total. If the system has a budget, the burn-down shows consumption against it and projects the overrun before it happens.
LLM Economics The efficiency view: cost per 1,000 requests and request volume per month, side by side. When you enable smart routing or the semantic cache in Govern, this is where the unit-cost line visibly bends.
Governance Guardrail blocks, policy denials, approvals — governance actually firing, not just configured.
Caching Cache hit rates and savings, response and prompt cache side by side.
Agents A2A traffic, agent authorization decisions, dry-run would-blocks from your grant rollout.
Reliability Breaker state, fallbacks, latency percentiles.

→ Mission Control

The system’s detail page (AI Estate → AI Systems → your system → Signals) is the per-system view of the same traffic, cut the way assurance cuts it — five sub-tabs: Liveness (the editable expectation and whether it is being met), Cost (cost per run and per completed task, with the declared bound from the contract beside it), Trajectory (turns and actions per run — a creeping loop shows up here first), Errors and Outcomes. Everything on it is computed from the run ledger, so it only fills once the run header is reaching the gateway.

The AI System’s group page (Resource Groups → your system → Usage) is the self-serve version — the view you check without leaving the group you’re configuring. Totals for requests, cost, tokens, MCP/agent/skill calls and error rate, activity charts that adapt their grain to the period (daily bars up to a month, weekly to a quarter, monthly beyond), and per-model / per-tool / per-capability breakdowns.

Every dashboard above meters requests. But you built an agent: one job is many requests. The run ledger — fed by the run header you wired in Develop — attributes cost to completed tasks:

“Refund handling costs $0.14 per resolved ticket” is a business number. “gpt-5.2 averaged 5,700 tokens per request” is not.

Runs live under AI Estate → AI Systems → your system → Runs: one row per task with its outcome, turns, actions, tools, cost, who closed it and its chain integrity; Details on any row opens everything the ledger holds about it, including the contract version it executed under. The Trajectory tab replays a run turn by turn. Cost-per-run belongs to Assure as a drift signal too — a creeping step count shows up there before it shows up on the invoice.

Logs — from symptom to evidence in one id

Section titled “Logs — from symptom to evidence in one id”

Every request through the gateway writes a proxy-log row: identity, group, model/tool, tokens, cost, latency, and the governance verdict (allowed / blocked / redacted / parked / denied) with the reason. Filter by your system’s group, or — the debugging move — send x-correlation-id on your agent’s requests and jump straight from a user complaint to the exact rows.

Terminal window
curl http://localhost:8100/v1/proxy/llm/chat/completions \
-H "Authorization: Bearer $KEY" -H "x-correlation-id: ticket-8841" ...

Governance events (who approved that refund, which admin changed a grant) land in the audit trail alongside. → Logs & audit trail

Two channels, split by what they ask of you:

  • Operational alerts — configure usage alerts on the system’s group (spend thresholds, error-rate spikes, volume anomalies). They land in the bell drawer on the Overview and in Slack or a webhook if you subscribe, with an acknowledgement workflow.
  • Findings — anything that needs a decision from a human: a liveness miss, a drift detection, a contract breach, a health transition, a failed run, a budget hard-stop, an expired approval. These go to Mission Control → Inbox, whose sidebar item carries a red badge for open critical items and a yellow one for warnings. Resolving one records a disposition and a note that the Assurance Report later cites. A liveness miss is raised in both places on purpose: paged from the alert, decided in the Inbox — so “the copilot went quiet on Saturday” is a notification, not a Monday discovery.

Your agent can also self-observe: governed responses carry X-Usage-Warning headers as budgets approach (Daily budget at 85%: $4.25 / $5.00) — back off before the hard 429.

The gateway exports OpenTelemetry — traces, metrics and logs into the Prometheus/Grafana/Loki stack you already run (Metrics & Grafana). And if part of your AI estate does not route through the gateway yet, two inlets bring it into view: Traffic Data Import pulls vendor admin telemetry in as observed traffic — visible in the same Mission Control views, clearly labeled as observed rather than governed — and the OpenTelemetry GenAI inlet accepts a system’s own trace so it becomes runs in the ledger, marked reported rather than witnessed on every row.

You can answer “what is it doing and what does it cost” in seconds, scoped to your system, at request and task granularity — and the system pages you when something needs you. The last stage makes the deeper promise: proving it behaves, and proving it to someone else.