Skip to content

Health & the Assurance Report

Everything assurance measures rolls up here: one health verdict per system, and one artifact that proves it.

health = min(liveness, behaviour, reliability, cost, conformance, oversight, review)

Worst-of, not an average. A system that is cheap, fast and dead is not 80% healthy — averaging is how a dashboard reports a comfortable number for a system nobody can call. Every component is shown alongside the total, because a score nobody can decompose is a score nobody trusts.

States: healthy · learning · degraded · drifting · stalled · silent · suspended · unknown.

Not every signal means something for every system

Section titled “Not every signal means something for every system”

The six kind-scoped components are not equally meaningful for every shape of AI System. Cadence detection is the headline signal for a scheduled integration and actively wrong for an assistant — humans are bursty and go on holiday, and paging someone about a quiet Sunday on a chat assistant is how an alert category gets muted. Trajectory drift is the headline signal for an agent and near-meaningless for an application, where the call shape is fixed by the product code.

So an AI System’s system_kind selects a signal profile, and each signal lands at one of three levels:

Level Scored? Counts toward the verdict? Rendered as
enforced Yes Yes The percentage
advisory Yes No The percentage, badged advisory
not_applicable No No n/a — never 0%, never a pass

not_applicable follows the platform’s standing rule that absence of evidence is never green: a signal that does not apply has not been satisfied, it was not asked. It never reaches the scorer as 1.0.

Kind Relaxed Why
agent — The strictest superset, and the default: a system we cannot classify is never scored leniently
workflow — Deterministic control flow is, if anything, more learnable than an agent
integration — Cadence is its headline signal
service liveness (advisory) Its callers decide when it runs, so silence describes them rather than it
application liveness, behaviour (advisory) Traffic follows product usage; call shape is fixed by product code, so a change is a deployment
assistant liveness (n/a), behaviour (advisory) A person may ask anything, and a quiet period is a weekend rather than an outage

Reliability, cost and conformance are never relaxed by any kind. They are shape-independent: an erroring, overspending or policy-violating system is unhealthy whether it is an agent or a chat box.

When liveness is not enforced for a kind, silence does not override the verdict either — otherwise the relaxation would be cosmetic and the quiet assistant would still read as an outage.

Oversight: was the human control effective?

Section titled “Oversight: was the human control effective?”

The other five components score the machine. Oversight scores the human control loop around it — because “a human approved it” and “a human exercised judgement” are different claims, and only the second one is oversight.

It applies only to systems that actually use approval gates. A system that requested no approvals in the window scores null, never zero: no gate configured means no oversight requirement was observed, which is a question about the system’s contract, not a failure of its oversight. null degrades nothing in the worst-of roll-up, exactly like every other component’s insufficient-evidence path.

What is scored

Measure Why it is scored
Lapse rate — approvals that expired with no human decision The one oversight failure with no benign reading: someone was asked and never answered
Approver concentration — every decision made by one person A control with no redundancy is a single point of failure. Applied as a ceiling, not a zero, and only once there are enough decisions for “one approver” to mean something

Incident windows excuse the rate, never the record. Oversight is a rate — holds decided over holds raised — so it is graded like the error rate: an approval hold raised inside a declared incident window of its AI System (the nearest AI System above the requester’s group), or a tenant-wide one, is not judged. The count is always stated beside the score — “N approval request(s) raised inside declared incident windows: recorded, not judged” — on the Signals overview, in the Assurance Report (the oversight section and Excluded from judgement) and on Compliance → Human Oversight, where requested / decided / lapsed / in flight stay the complete record. If every hold in the period was raised inside a window, oversight is unmeasured (null), not passed.

A window only excuses what it covers. A hold raised outside every window that lapses because nobody decided it in time — overnight, say — is a configuration or operations matter and always counts: set an approval window long enough for the people on call. There is no other excuse for a lapsed hold — no grace period, no automatic pardon. A hold is judged by when it was raised, not when it lapsed: one raised just before a window opened counts, even if it lapsed inside it.

What is deliberately not scored

Decision latency and denial rate are recorded, published in the report, and trended — but never thresholded. A fast approval may be an alert, well-staffed team, and a genuinely well-behaved system is legitimately approved every time. There is no honest absolute bound for either, and inventing one would manufacture findings.

Instead both are learned as baselines, so movement is caught by the ordinary drift machinery with a cause attached — “this team used to deliberate for four minutes and now answers in six seconds” is detectable and defensible in a way that a hardcoded limit is not. Rubber-stamping surfaces as drift, not as a rule.

Review — the human loop around the findings

Section titled “Review — the human loop around the findings”

The seventh component (RFC 0017) closes the loop on the Assurance Inbox: when the platform flags something for human review — a guardrail block, drift, a contract breach — somebody accountable has to look. An open critical inbox item older than the review SLA (7 days by default) scores review at 0.5, which the worst-of roll-up turns into degraded: a system whose flagged findings nobody reads is not healthy, whatever its other numbers say.

null until the system has ever had an inbox item — never having been flagged is not evidence of review, and unmeasured is not 100%. Review is enforced for every system kind: unreviewed findings are a human-process failure regardless of the system’s shape.

intended_use and the intended clients are declarations. Intended clients come in two kinds: client labels (intended_client_labels, the apps and channels that call the system) and client systems (intended_client_systems, other AI Systems allowed to delegate to it). Conformance compares them against what the gateway actually saw, using the run’s client label for the first and, for the second, the immediate calling system of each run segment — in A → B → C, C’s caller is B (the system that declared the dependency and made the call), not A, whose run it is.

A client label names a channel, never an HTTP library: it is the X-Brutor-Client header when the caller sends one, else the product name of the first User-Agent token without its version (python-httpx/0.28.1 → python-httpx, a browser → browser). And a system is never its own client: runs made by one of the system’s own member agent identities are not judged against intended_client_labels — they are the system at work, counted as observed traffic and judged on implementation and release instead. A system whose only caller is its own agent needs no client labels at all.

What conformance reports:

  • An undeclared client that called the system — on a high-risk system this is escalated, because it is the finding a risk owner reads first.
  • A declared client that never called — reported only once there is enough traffic to distinguish “absent” from “quiet this week”.
  • A granted resource never used — a least-privilege finding: blast radius nobody is getting value from.

A caller system that is not in intended_client_systems is the finding composition.unaccepted_caller. By default it is advisory — the call is served and the finding says so. A service can opt into enforcement per system with enforce_accepted_callers: true (AI Systems only; a signed contract term, so turning it on or off shows in the next contract diff as narrowed or widened): the gateway then refuses an A2A call from an unaccepted caller before the call reaches the service (HTTP 403, proxy subtype unaccepted_caller, a sealed denied record), and the finding counts the refused attempts and says enforced.

A system with no traffic has no conformance score. It has not demonstrated conformance; it has demonstrated nothing.

Health says now; the report’s verdict says over the period, worst-first:

Status Means
violations At least one signed envelope term was breached
insufficient_evidence No runs, or too few to evaluate any declared term — never rounded up to assured
assured_with_exceptions Within its envelope, with governed refusals, open drift or unevaluable terms on record
assured “The system remained within its approved operating envelope during the reporting period.”

A violation observed through thin evidence is still a violation — thin evidence weakens an assured claim, never a breach. The verdict sits above a per-term contract compliance table, an exceptions section (guardrail blocks, policy denials, approval escalations, automatic responses — counted, never netted away) and a changes section (contract versions, attributed model and dependency changes, current config drift).

Every verdict also states its basis. When an approved operating envelope is in force, the verdict is measured against the signature (signed_envelope). When nothing has been signed off, the verdict is conditional: it is measured against the system’s own observed history (observed_history), says so in its conclusion, and never borrows the authority of a signature that does not exist. “Assured” over an unsigned envelope would be a contradiction a customer notices immediately — so the report doesn’t say it. A third basis, reported_telemetry, applies when every run in the period was ingested rather than proxied: the gateway witnessed none of them and enforced nothing, so even a signed envelope cannot lift the verdict above conditional.

In the console the Assurance tab opens on Conclusion — the status, the executive rows, the verdict and contract compliance, headed by a diagram of the assurance loop with the node this system has reached marked and the number each node rests on (runs in the ledger, baselines ready, health, verdict, findings decided, snapshots captured). Evidence holds the health components, the maturity ladder, the baselines and what the report does not cover; Exceptions & changes the exceptions, changes and recommended controls; Snapshots the captured certificates.

One artifact per AI System answering all four guarantees with evidence: identity and declared intent, contract history with approvers and hashes, the learned baselines, drift events with their causes and the responses taken, run outcomes and cost per completed run, liveness record, replay results, conformance, and recommended controls.

The engineer opens it at 09:00; the auditor exports it in December. The same artifact serves both.

The Assurance tab’s Conclusion sub-tab: the assurance loop with this system at A human decides and its open findings as a button; the status Assured with exceptions · operating envelope with its basis note; the executive rows — what we know, what prevents full assurance, recommended action, why this matters; the five dimensions with their reasons; the period verdict; and the contract compliance table with every term met.

The report opens with a plain-language assurance status — the word a decision-maker reads first, derived from the verdict and its basis:

Status Means
Assured · operating envelope Within its approved envelope, verdict measured against a signature
Assured with exceptions · operating envelope Within its signed envelope, with governed exceptions on record
Conditionally assured · observed history Within observed historical behaviour only — no approved envelope is in force, or the period’s evidence was entirely reported telemetry
Not assured — contract violations A signed envelope term was breached
Evidence insufficient Too little evidence to support any conclusion

The label carries its basis on its face, and a one-line basis_note under the badge says what the claim was measured against. Under the status sit three labelled rows, synthesised deterministically from facts stated elsewhere in the report (the summary never knows something the sections below cannot prove): What we know — what the evidence supports; What prevents full assurance — the named deficiencies (an unapproved contract, unsigned execution chains, unexplained drift, missing independent evidence); Recommended action — rely on it, remain under monitored operation, or treat it as out of bounds; and Why this matters — the risk interpretation in plain terms (“No evidence of abnormal behaviour was detected … however, it cannot yet be established that the system remains within an approved operating boundary”), so a busy reader never has to translate a list of missing evidence into “is this system okay or not?” themselves.

Beside the conclusion sit five assurance dimensions, each ok, partial or missing with its reason: operational evidence, behavioural stability, contract assurance, execution evidence, and independent evidence. One merged verdict hides that operational evidence can be complete while contract assurance has not even started; the dimensions show which half of assurance is done.

Values below are from an actual export (spec 1.4, 720-hour period) for a system called Contract Review Assistant, owned by Legal & Compliance, kind assistant, risk tier minimal. It is a useful example precisely because it is not a clean bill of health.

Status: Conditionally assured

Row What the report says
What we know Operated within its observed historical behaviour during the reporting period, but formal assurance is incomplete
What prevents full assurance The operating contract is not approved · 40% of runs lack cryptographically signed execution chains · no independent evidence is attached
Recommended action Remain under monitored operation until the outstanding assurance requirements are completed
Why this matters No evidence of abnormal behaviour was detected. However, it cannot yet be established that the system remains within an approved operating boundary — behaving like itself is not the same as behaving as approved

The five dimensions, and why the maturity is 0

Dimension Status Reason given
Operational evidence ok 99 runs fully observed in the period
Behavioural stability missing Baselines are still learning; behaviour cannot be judged against normal yet
Contract assurance missing No approved operating envelope is in force
Execution evidence partial Only 60% of runs carry gateway-signed chains; no replay evidence
Independent evidence missing No external evidence is attached

Because the ladder is cumulative, one missing on behavioural stability pins the whole system at Evidence maturity 0 — Observable, with 1 — Baseline established named as the next rung. The verdict alongside reads assured, qualified conditional — measured against observed history: a real claim about a real period, scoped honestly to what it was measured against.

Learned baselines are published with their own state, so a reader can see which numbers are load-bearing yet:

action_count 3 ± 1 learning · 46 runs
cost_per_run 0.0156115 ± 0.004335 learning · 46 runs
model_mix gpt-5.5 100% learning · 46 runs
terminal_state_mix completed 78% · abandoned 9% learning · 46 runs
· completed_degraded 9%
tool_error_rate 0 ± 0 learning · 36 runs

Exceptions in the period were all zero — including Approval requests: 0, which is exactly the case where oversight scores null rather than zero: this system exercised no approval gate, so there was no human control to measure. That is a fact about its contract, not a mark against it.

The report closes with recommended controls ranked by severity — here bind_replay_suite, mint_contract and sign_run_chains at medium, revoke_unused_grants at low — each naming the deficiency it would close.

The Evidence sub-tab: the worst-of health block — Silent, 50%, with the seven components and oversight not measured; the evidence maturity ladder at level 2, Controlled, with what blocks level 3 and what to do about it; and the learned baselines, each with its value, its run count and whether it is ready or still learning.

Deliberately named evidence maturity, not “assurance level”: the number measures how much evidence exists, never how trustworthy the system is — a system at maturity 0 with a clean period is conditionally assured, not “basically unassured”, and the status above carries that claim. For the same reason the executive area shows only the word (“Evidence maturity: Observable”); the numeric rung and the path to the next one live in the detail sections, where numbers belong. The criteria:

Level Name Reached when
0 Observable Traffic and behaviour are visible through the gateway
1 Baseline established The system’s normal operating behaviour has been learned
2 Controlled An approved contract binds the system; enforcement is active
3 Evidence-backed Execution chains are signed and replay evidence exists
4 Independently assured Independent external evidence is attached and unexpired

The ladder is cumulative — a level is only reached when every rung below it holds, so signed chains never lift a system past a missing contract.

It is rendered as a checklist, not a score. Every rung is listed with its requirement and whether it is held, including rungs held above a broken one: signed execution chains do not raise the level while the contract is missing, but hiding that the work is already done is how an operator ends up doing it twice.

Alongside it, the report names the next rung’s blocker and the single action that closes it:

To reach L2 Controlled
Blocking now No approved operating envelope is in force
Do this Declare an operating envelope and have it signed off

The blocker text is read from the evaluated dimension itself rather than written a second time, so it cannot drift from the logic that decided the rung was unmet — a ladder that confidently points at the wrong thing is worse than one that says nothing. This lines up with the recommended controls the report already proposes.

Drift findings in the report are printed as observable facts, not labels: each carries the baseline value, the observed value, the direction of movement and the effect size — so “3 open drift findings” reads as tool calls per run moved from 2 to 4.7, up, high severity, unexplained, answerable from the report alone.

The baselines section prints the yardstick the drift section is measured against — what normal looks like for this system, learned from its own run ledger (median/MAD; run metrics are heavy-tailed), per metric with its learning state and sample count. Seasonal and per-actor sub-baselines are counted, not dumped: they exist to sharpen detection, and a hundred day-of-week rows would bury the signal in an auditor-facing document.

The recommended controls section turns the report’s findings into named next controls — an unexplained open drift finding suggests a response policy, an undeclared client suggests declaring or blocking it, a declared-but-unsigned envelope suggests approving the contract, a high-risk system without an unexpired impact assessment is flagged high severity. Strictly rule-based over facts already in the report, worst first, advisory: proposing a control and enacting one are different authorities, and a fully governed system’s list is empty rather than padded.

The report also lists the system’s external evidence — impact assessments, evaluation, red-team and pen-test reports — as verifiable pointers (uri + sha256, assessor, independence, validity). When nothing is attached, or nothing attached is independently produced, the report says so in not_covered rather than letting platform-generated evidence pass as the whole story.

The Composition and dependencies section (spec 1.6; 1.7 adds member agents) lists every declared dependency of the system: its role (shared service or external), criticality, whether it contributes to the decision (with the sealed separability rationale when it does not), its current status — the same dependency-health entry the health endpoint shows — the evidence basis (tested for a shared service assured from its segments, assumed for an external agent seen only at the boundary, with its A2A task-state mix), the contract version it runs against (followed or pinned), its supplier and whether the supplier’s written agreement is on file, plus the system’s callers, its declared and effective risk tier, and every open composition finding. Since spec 1.7 it also lists each member agent’s current release (name@version, build), its trust (verified or declared) and whether it is approved, with the open agent findings — so an unapproved release or an undeclared implementation change is sealed with the report. It is part of the frozen report, so the snapshot’s digest and signature cover it.

The Judgement exclusions section (spec 1.8) lists every incident window overlapping the period — who declared it, when, why, its cause category and how many of this system’s runs it excluded — and every baseline epoch with the resolution that wrote it. The runs inside a window are graded by nothing in the report (rates, envelope terms, conformance, the verdict), yet they are still in the ledger, sealed and exportable, and the run facts disclose how many were excluded. The coverage block states the excluded period in hours and runs — that period is unmeasured, not passed — and drift events carry the kind they were closed with.

A call’s guardrail evidence is not “checked” or “unchecked”. The resilience ladder means a check can have run on the configured detector, on a retry, or on the local built-in after the remote provider went down — and it can have been skipped entirely.

Tier Content inspected? Counts toward guardrail coverage
primary Yes, by the configured detector Yes
retried Yes, after a transient failure Yes
fallback Yes, by the local built-in — lower fidelity Yes, disclosed as degraded
break_glass No — an operator declared an outage Never
policy No — nothing answered; fail-open let it through Never

The first three are recorded in proxy_logs.guardrails_degraded, the last two in guardrails_skipped. Keeping them in separate columns is what lets the report say “this check ran, on the backup detector” without either inflating that into a full pass or discarding a real inspection as no inspection at all.

Windows where an operator declared a break-glass are named in not_covered with the reason and the period, because a report that quietly omits the hour when checks were off is the kind an auditor stops trusting.

Every section above answers did the traffic we saw behave? The report also answers the prior question — did we see the traffic? That is the gap an auditor opens with: a report can be immaculate about 340,000 governed calls and still be worthless if the team runs an ungoverned side channel beside them.

The coverage section states what fraction of the discovered AI estate is actually routed, with the basis the number rests on:

Basis Meaning
measured A collector reported recently; the ratio is current
stale Collectors exist but have gone quiet — the estate as last observed, not as it is now
unmeasured No collector has ever reported. No percentage is emitted at all

Two further honesty rules:

  • Unreconciled assets count against coverage. An asset the platform could not match to a configured route is not a routed asset, so folding it into the numerator would let a flaky reconcile pass inflate the headline.
  • Dismissals are disclosed, not hidden. An operator can raise the percentage by dismissing assets — a legitimate action on a real estate — so the dismissed count is reported alongside the ratio rather than silently subtracted from it.

Coverage also states how each observed run was seen. Runs that reached the ledger through the OpenTelemetry inlet count as observed — they are complete as reported — but their share is disclosed as external_pct, and a report whose evidence is entirely reported carries the reported_telemetry basis above. Assurance covers what the gateway witnessed plus what was ingested with its provenance stated, never the whole estate by assumption.

The section also flags planes with no visibility at all. The commonest real gap is a fully-instrumented network with zero endpoint coverage: laptops running local models that no attestation can see until Scout is deployed.

Snapshots are immutable. Capturing one freezes the JSON; recomputing it later would answer “what is true now” rather than “what was true then”, and evidence that changes when the system does is not evidence.

Snapshots are also attested: at capture time the report’s SHA-256 manifest is signed with the deployment audit key — the same Ed25519 key that signs audit-chain checkpoints. Downloading a snapshot yields a portable AI System Assurance Certificate (spec: ai.brutor/ai-system-assurance-report, schema at GET /v1/admin/standards/ai-system-assurance-report/schema) that a recipient verifies offline with sha256sum and any Ed25519 library — no API call back to the platform. Signing is best-effort and honest: a deployment without an audit key produces certificates that say signed: false with the reason, never an implied attestation.

Set assurance_capture_interval_hours on the tenant to capture automatically — continuous post-market monitoring evidence produced by a machine rather than by someone remembering to click. It is off by default: a job nobody switched on should not silently accumulate a snapshot of every system every day.

Captured reports and ready baselines feed the AIMS control coverage automatically: A.6.2.4 (performance evaluation — behavioural baselines exist, so “performing as expected” is measurable) and A.6.2.7 (post-market monitoring — dated evidence keeps being produced). Obligation evidence for ISO/IEC 42001 and every other framework — sealed, witnessed and stated n of m — lives in the Compliance section and its Evidence Reports.