Health & the Assurance Report
Everything assurance measures rolls up here: one health verdict per system, and one artifact that proves it.
One number, seven components
Section titled “One number, seven components”health = min(liveness, behaviour, reliability, cost, conformance, oversight, review)Worst-of, not an average. A system that is cheap, fast and dead is not 80% healthy — averaging is how a dashboard reports a comfortable number for a system nobody can call. Every component is shown alongside the total, because a score nobody can decompose is a score nobody trusts.
States: healthy · learning · degraded · drifting · stalled · silent ·
suspended · unknown.
Not every signal means something for every system
Section titled “Not every signal means something for every system”The six kind-scoped components are not equally meaningful for every shape of AI System. Cadence
detection is the headline signal for a scheduled integration and actively wrong for an
assistant — humans are bursty and go on holiday, and paging someone about a quiet Sunday
on a chat assistant is how an alert category gets muted. Trajectory drift is the headline
signal for an agent and near-meaningless for an application, where the call shape is
fixed by the product code.
So an AI System’s system_kind selects a signal profile, and each
signal lands at one of three levels:
| Level | Scored? | Counts toward the verdict? | Rendered as |
|---|---|---|---|
enforced |
Yes | Yes | The percentage |
advisory |
Yes | No | The percentage, badged advisory |
not_applicable |
No | No | n/a — never 0%, never a pass |
not_applicable follows the platform’s standing rule that absence of evidence is never
green: a signal that does not apply has not been satisfied, it was not asked. It never
reaches the scorer as 1.0.
| Kind | Relaxed | Why |
|---|---|---|
agent |
— | The strictest superset, and the default: a system we cannot classify is never scored leniently |
workflow |
— | Deterministic control flow is, if anything, more learnable than an agent |
integration |
— | Cadence is its headline signal |
service |
liveness (advisory) | Its callers decide when it runs, so silence describes them rather than it |
application |
liveness, behaviour (advisory) | Traffic follows product usage; call shape is fixed by product code, so a change is a deployment |
assistant |
liveness (n/a), behaviour (advisory) | A person may ask anything, and a quiet period is a weekend rather than an outage |
Reliability, cost and conformance are never relaxed by any kind. They are shape-independent: an erroring, overspending or policy-violating system is unhealthy whether it is an agent or a chat box.
When liveness is not enforced for a kind, silence does not override the verdict either — otherwise the relaxation would be cosmetic and the quiet assistant would still read as an outage.
Oversight: was the human control effective?
Section titled “Oversight: was the human control effective?”The other five components score the machine. Oversight scores the human control loop around it — because “a human approved it” and “a human exercised judgement” are different claims, and only the second one is oversight.
It applies only to systems that actually use approval gates. A system that requested no
approvals in the window scores null, never zero: no gate configured means no
oversight requirement was observed, which is a question about the system’s contract,
not a failure of its oversight. null degrades nothing in the worst-of roll-up, exactly
like every other component’s insufficient-evidence path.
What is scored
| Measure | Why it is scored |
|---|---|
| Lapse rate — approvals that expired with no human decision | The one oversight failure with no benign reading: someone was asked and never answered |
| Approver concentration — every decision made by one person | A control with no redundancy is a single point of failure. Applied as a ceiling, not a zero, and only once there are enough decisions for “one approver” to mean something |
Incident windows excuse the rate, never the record. Oversight is a rate — holds decided
over holds raised — so it is graded like the error rate: an approval hold raised inside a
declared incident window of its AI System (the nearest
AI System above the requester’s group), or a tenant-wide one, is not judged. The count is
always stated beside the score — “N approval request(s) raised inside declared incident
windows: recorded, not judged” — on the Signals overview, in the Assurance Report (the
oversight section and Excluded from judgement) and on Compliance → Human Oversight, where
requested / decided / lapsed / in flight stay the complete record. If every hold in the
period was raised inside a window, oversight is unmeasured (null), not passed.
A window only excuses what it covers. A hold raised outside every window that lapses because nobody decided it in time — overnight, say — is a configuration or operations matter and always counts: set an approval window long enough for the people on call. There is no other excuse for a lapsed hold — no grace period, no automatic pardon. A hold is judged by when it was raised, not when it lapsed: one raised just before a window opened counts, even if it lapsed inside it.
What is deliberately not scored
Decision latency and denial rate are recorded, published in the report, and trended — but never thresholded. A fast approval may be an alert, well-staffed team, and a genuinely well-behaved system is legitimately approved every time. There is no honest absolute bound for either, and inventing one would manufacture findings.
Instead both are learned as baselines, so movement is caught by the ordinary drift machinery with a cause attached — “this team used to deliberate for four minutes and now answers in six seconds” is detectable and defensible in a way that a hardcoded limit is not. Rubber-stamping surfaces as drift, not as a rule.
Review — the human loop around the findings
Section titled “Review — the human loop around the findings”The seventh component (RFC 0017) closes the loop on the Assurance
Inbox: when the platform
flags something for human review — a guardrail block, drift, a contract
breach — somebody accountable has to look. An open critical inbox item
older than the review SLA (7 days by default) scores review at 0.5,
which the worst-of roll-up turns into degraded: a system whose flagged
findings nobody reads is not healthy, whatever its other numbers say.
null until the system has ever had an inbox item — never having been
flagged is not evidence of review, and unmeasured is not 100%. Review is
enforced for every system kind: unreviewed findings are a human-process
failure regardless of the system’s shape.
Conformance: declared vs observed
Section titled “Conformance: declared vs observed”intended_use and the intended clients are declarations. Intended clients come in two
kinds: client labels (intended_client_labels, the apps and channels that call the
system) and client systems (intended_client_systems, other AI Systems allowed to
delegate to it). Conformance compares them against what the gateway actually saw, using
the run’s client label for the first and, for the second, the immediate calling system
of each run segment — in A → B → C, C’s caller is B (the system that declared the
dependency and made the call), not A, whose run it is.
A client label names a channel, never an HTTP library: it is the X-Brutor-Client
header when the caller sends one, else the product name of the first User-Agent token
without its version (python-httpx/0.28.1 → python-httpx, a browser → browser). And a
system is never its own client: runs made by one of the system’s own member agent
identities are not judged against intended_client_labels — they are the system at work,
counted as observed traffic and judged on
implementation and release instead. A system whose
only caller is its own agent needs no client labels at all.
What conformance reports:
- An undeclared client that called the system — on a high-risk system this is escalated, because it is the finding a risk owner reads first.
- A declared client that never called — reported only once there is enough traffic to distinguish “absent” from “quiet this week”.
- A granted resource never used — a least-privilege finding: blast radius nobody is getting value from.
A caller system that is not in intended_client_systems is the finding
composition.unaccepted_caller. By default it is advisory — the call is served and the
finding says so. A service can opt into enforcement per system with
enforce_accepted_callers: true (AI Systems only; a signed contract term, so turning it on
or off shows in the next contract diff as narrowed or widened): the gateway then refuses an
A2A call from an unaccepted caller before the call reaches the service (HTTP 403, proxy
subtype unaccepted_caller, a sealed denied record), and the finding counts the refused
attempts and says enforced.
A system with no traffic has no conformance score. It has not demonstrated conformance; it has demonstrated nothing.
The period verdict
Section titled “The period verdict”Health says now; the report’s verdict says over the period, worst-first:
| Status | Means |
|---|---|
violations |
At least one signed envelope term was breached |
insufficient_evidence |
No runs, or too few to evaluate any declared term — never rounded up to assured |
assured_with_exceptions |
Within its envelope, with governed refusals, open drift or unevaluable terms on record |
assured |
“The system remained within its approved operating envelope during the reporting period.” |
A violation observed through thin evidence is still a violation — thin evidence weakens an assured claim, never a breach. The verdict sits above a per-term contract compliance table, an exceptions section (guardrail blocks, policy denials, approval escalations, automatic responses — counted, never netted away) and a changes section (contract versions, attributed model and dependency changes, current config drift).
Every verdict also states its basis. When an approved operating envelope is in
force, the verdict is measured against the signature (signed_envelope). When nothing
has been signed off, the verdict is conditional: it is measured against the
system’s own observed history (observed_history), says so in its conclusion, and
never borrows the authority of a signature that does not exist. “Assured” over an
unsigned envelope would be a contradiction a customer notices immediately — so the
report doesn’t say it. A third basis, reported_telemetry, applies when every run in
the period was ingested rather than proxied: the gateway
witnessed none of them and enforced nothing, so even a signed envelope cannot lift the
verdict above conditional.
The Assurance Report
Section titled “The Assurance Report”In the console the Assurance tab opens on Conclusion — the status, the executive rows, the verdict and contract compliance, headed by a diagram of the assurance loop with the node this system has reached marked and the number each node rests on (runs in the ledger, baselines ready, health, verdict, findings decided, snapshots captured). Evidence holds the health components, the maturity ladder, the baselines and what the report does not cover; Exceptions & changes the exceptions, changes and recommended controls; Snapshots the captured certificates.
One artifact per AI System answering all four guarantees with evidence: identity and declared intent, contract history with approvers and hashes, the learned baselines, drift events with their causes and the responses taken, run outcomes and cost per completed run, liveness record, replay results, conformance, and recommended controls.
The engineer opens it at 09:00; the auditor exports it in December. The same artifact serves both.
The first twenty seconds
Section titled “The first twenty seconds”
The report opens with a plain-language assurance status — the word a decision-maker reads first, derived from the verdict and its basis:
| Status | Means |
|---|---|
| Assured · operating envelope | Within its approved envelope, verdict measured against a signature |
| Assured with exceptions · operating envelope | Within its signed envelope, with governed exceptions on record |
| Conditionally assured · observed history | Within observed historical behaviour only — no approved envelope is in force, or the period’s evidence was entirely reported telemetry |
| Not assured — contract violations | A signed envelope term was breached |
| Evidence insufficient | Too little evidence to support any conclusion |
The label carries its basis on its face, and a one-line basis_note under the badge says
what the claim was measured against. Under the status sit three labelled rows, synthesised deterministically from facts stated
elsewhere in the report (the summary never knows something the sections below cannot
prove): What we know — what the evidence supports; What prevents full assurance —
the named deficiencies (an unapproved contract, unsigned execution chains, unexplained
drift, missing independent evidence); Recommended action — rely on it, remain
under monitored operation, or treat it as out of bounds; and Why this matters — the
risk interpretation in plain terms (“No evidence of abnormal behaviour was detected …
however, it cannot yet be established that the system remains within an approved
operating boundary”), so a busy reader never has to translate a list of missing evidence
into “is this system okay or not?” themselves.
Beside the conclusion sit five assurance dimensions, each ok, partial or missing
with its reason: operational evidence, behavioural stability, contract assurance,
execution evidence, and independent evidence. One merged verdict hides that operational
evidence can be complete while contract assurance has not even started; the dimensions
show which half of assurance is done.
A real report, read end to end
Section titled “A real report, read end to end”Values below are from an actual export (spec 1.4, 720-hour period) for a system called
Contract Review Assistant, owned by Legal & Compliance, kind assistant, risk tier
minimal. It is a useful example precisely because it is not a clean bill of health.
Status: Conditionally assured
| Row | What the report says |
|---|---|
| What we know | Operated within its observed historical behaviour during the reporting period, but formal assurance is incomplete |
| What prevents full assurance | The operating contract is not approved · 40% of runs lack cryptographically signed execution chains · no independent evidence is attached |
| Recommended action | Remain under monitored operation until the outstanding assurance requirements are completed |
| Why this matters | No evidence of abnormal behaviour was detected. However, it cannot yet be established that the system remains within an approved operating boundary — behaving like itself is not the same as behaving as approved |
The five dimensions, and why the maturity is 0
| Dimension | Status | Reason given |
|---|---|---|
| Operational evidence | ok |
99 runs fully observed in the period |
| Behavioural stability | missing |
Baselines are still learning; behaviour cannot be judged against normal yet |
| Contract assurance | missing |
No approved operating envelope is in force |
| Execution evidence | partial |
Only 60% of runs carry gateway-signed chains; no replay evidence |
| Independent evidence | missing |
No external evidence is attached |
Because the ladder is cumulative, one missing on behavioural stability pins the whole
system at Evidence maturity 0 — Observable, with 1 — Baseline established named as
the next rung. The verdict alongside reads assured, qualified conditional — measured
against observed history: a real claim about a real period, scoped honestly to what it
was measured against.
Learned baselines are published with their own state, so a reader can see which numbers are load-bearing yet:
action_count 3 ± 1 learning · 46 runscost_per_run 0.0156115 ± 0.004335 learning · 46 runsmodel_mix gpt-5.5 100% learning · 46 runsterminal_state_mix completed 78% · abandoned 9% learning · 46 runs · completed_degraded 9%tool_error_rate 0 ± 0 learning · 36 runsExceptions in the period were all zero — including Approval requests: 0, which is
exactly the case where oversight scores
null rather than zero: this system exercised no approval gate, so there was no human
control to measure. That is a fact about its contract, not a mark against it.
The report closes with recommended controls ranked by severity — here
bind_replay_suite, mint_contract and sign_run_chains at medium, revoke_unused_grants
at low — each naming the deficiency it would close.
The evidence maturity ladder
Section titled “The evidence maturity ladder”
Deliberately named evidence maturity, not “assurance level”: the number measures how much evidence exists, never how trustworthy the system is — a system at maturity 0 with a clean period is conditionally assured, not “basically unassured”, and the status above carries that claim. For the same reason the executive area shows only the word (“Evidence maturity: Observable”); the numeric rung and the path to the next one live in the detail sections, where numbers belong. The criteria:
| Level | Name | Reached when |
|---|---|---|
| 0 | Observable | Traffic and behaviour are visible through the gateway |
| 1 | Baseline established | The system’s normal operating behaviour has been learned |
| 2 | Controlled | An approved contract binds the system; enforcement is active |
| 3 | Evidence-backed | Execution chains are signed and replay evidence exists |
| 4 | Independently assured | Independent external evidence is attached and unexpired |
The ladder is cumulative — a level is only reached when every rung below it holds, so signed chains never lift a system past a missing contract.
It is rendered as a checklist, not a score. Every rung is listed with its requirement and whether it is held, including rungs held above a broken one: signed execution chains do not raise the level while the contract is missing, but hiding that the work is already done is how an operator ends up doing it twice.
Alongside it, the report names the next rung’s blocker and the single action that closes it:
| To reach | L2 Controlled |
| Blocking now | No approved operating envelope is in force |
| Do this | Declare an operating envelope and have it signed off |
The blocker text is read from the evaluated dimension itself rather than written a second time, so it cannot drift from the logic that decided the rung was unmet — a ladder that confidently points at the wrong thing is worse than one that says nothing. This lines up with the recommended controls the report already proposes.
Drift findings in the report are printed as observable facts, not labels: each carries the baseline value, the observed value, the direction of movement and the effect size — so “3 open drift findings” reads as tool calls per run moved from 2 to 4.7, up, high severity, unexplained, answerable from the report alone.
The baselines section prints the yardstick the drift section is measured against — what normal looks like for this system, learned from its own run ledger (median/MAD; run metrics are heavy-tailed), per metric with its learning state and sample count. Seasonal and per-actor sub-baselines are counted, not dumped: they exist to sharpen detection, and a hundred day-of-week rows would bury the signal in an auditor-facing document.
The recommended controls section turns the report’s findings into named next controls — an unexplained open drift finding suggests a response policy, an undeclared client suggests declaring or blocking it, a declared-but-unsigned envelope suggests approving the contract, a high-risk system without an unexpired impact assessment is flagged high severity. Strictly rule-based over facts already in the report, worst first, advisory: proposing a control and enacting one are different authorities, and a fully governed system’s list is empty rather than padded.
The report also lists the system’s external evidence
— impact assessments, evaluation, red-team and pen-test reports — as verifiable pointers
(uri + sha256, assessor, independence, validity). When nothing is attached, or nothing
attached is independently produced, the report says so in not_covered rather than
letting platform-generated evidence pass as the whole story.
The Composition and dependencies section (spec 1.6; 1.7 adds member agents) lists every declared dependency
of the system: its role (shared service or external), criticality, whether it contributes
to the decision (with the sealed separability rationale when it does not), its current
status — the same dependency-health entry the health endpoint shows — the evidence basis
(tested for a shared service assured from its segments, assumed for an external agent
seen only at the boundary, with its A2A task-state mix), the contract version it runs
against (followed or pinned), its supplier and whether the supplier’s written agreement is
on file, plus the system’s callers, its declared and effective risk tier, and every open
composition finding. Since spec 1.7 it also lists each member agent’s current release
(name@version, build), its trust (verified or declared) and whether it is approved,
with the open agent findings — so an
unapproved release or an undeclared implementation change is sealed with the report. It is
part of the frozen report, so the snapshot’s digest and signature cover it.
The Judgement exclusions section (spec 1.8) lists every incident window overlapping the period — who declared it, when, why, its cause category and how many of this system’s runs it excluded — and every baseline epoch with the resolution that wrote it. The runs inside a window are graded by nothing in the report (rates, envelope terms, conformance, the verdict), yet they are still in the ledger, sealed and exportable, and the run facts disclose how many were excluded. The coverage block states the excluded period in hours and runs — that period is unmeasured, not passed — and drift events carry the kind they were closed with.
Guardrail coverage is tiered, not binary
Section titled “Guardrail coverage is tiered, not binary”A call’s guardrail evidence is not “checked” or “unchecked”. The resilience ladder means a check can have run on the configured detector, on a retry, or on the local built-in after the remote provider went down — and it can have been skipped entirely.
| Tier | Content inspected? | Counts toward guardrail coverage |
|---|---|---|
primary |
Yes, by the configured detector | Yes |
retried |
Yes, after a transient failure | Yes |
fallback |
Yes, by the local built-in — lower fidelity | Yes, disclosed as degraded |
break_glass |
No — an operator declared an outage | Never |
policy |
No — nothing answered; fail-open let it through | Never |
The first three are recorded in proxy_logs.guardrails_degraded, the last two in
guardrails_skipped. Keeping them in separate columns is what lets the report say “this
check ran, on the backup detector” without either inflating that into a full pass or
discarding a real inspection as no inspection at all.
Windows where an operator declared a break-glass are named in not_covered with the
reason and the period, because a report that quietly omits the hour when checks were off
is the kind an auditor stops trusting.
Coverage attestation
Section titled “Coverage attestation”Every section above answers did the traffic we saw behave? The report also answers the prior question — did we see the traffic? That is the gap an auditor opens with: a report can be immaculate about 340,000 governed calls and still be worthless if the team runs an ungoverned side channel beside them.
The coverage section states what fraction of the
discovered AI estate is actually routed, with the
basis the number rests on:
| Basis | Meaning |
|---|---|
measured |
A collector reported recently; the ratio is current |
stale |
Collectors exist but have gone quiet — the estate as last observed, not as it is now |
unmeasured |
No collector has ever reported. No percentage is emitted at all |
Two further honesty rules:
- Unreconciled assets count against coverage. An asset the platform could not match to a configured route is not a routed asset, so folding it into the numerator would let a flaky reconcile pass inflate the headline.
- Dismissals are disclosed, not hidden. An operator can raise the percentage by dismissing assets — a legitimate action on a real estate — so the dismissed count is reported alongside the ratio rather than silently subtracted from it.
Coverage also states how each observed run was seen. Runs that reached the ledger through
the OpenTelemetry inlet count as observed — they are
complete as reported — but their share is disclosed as external_pct, and a report
whose evidence is entirely reported carries the reported_telemetry basis above.
Assurance covers what the gateway witnessed plus what was ingested with its provenance
stated, never the whole estate by assumption.
The section also flags planes with no visibility at all. The commonest real gap is a fully-instrumented network with zero endpoint coverage: laptops running local models that no attestation can see until Scout is deployed.
Snapshots are immutable. Capturing one freezes the JSON; recomputing it later would answer “what is true now” rather than “what was true then”, and evidence that changes when the system does is not evidence.
Snapshots are also attested: at capture time the report’s SHA-256 manifest is signed
with the deployment audit key — the same Ed25519 key that signs
audit-chain checkpoints.
Downloading a snapshot yields a portable AI System Assurance Certificate
(spec: ai.brutor/ai-system-assurance-report, schema at
GET /v1/admin/standards/ai-system-assurance-report/schema) that a recipient verifies
offline with sha256sum and any Ed25519 library — no API call back to the platform.
Signing is best-effort and honest: a deployment without an audit key produces
certificates that say signed: false with the reason, never an implied attestation.
Set assurance_capture_interval_hours on the tenant to capture automatically —
continuous post-market monitoring evidence produced by a machine rather than by someone
remembering to click. It is off by default: a job nobody switched on should not silently
accumulate a snapshot of every system every day.
ISO 42001
Section titled “ISO 42001”Captured reports and ready baselines feed the AIMS control coverage automatically:
A.6.2.4 (performance evaluation — behavioural baselines exist, so “performing as
expected” is measurable) and A.6.2.7 (post-market monitoring — dated evidence keeps
being produced). Obligation evidence for ISO/IEC 42001 and every other framework — sealed, witnessed and stated n of m — lives in the Compliance section and its Evidence Reports.
Related
Section titled “Related”- How assurance works — the lifecycle this page is the verdict of
- Assurance API — health, reports and capture endpoints
- Mission Control — the fleet board the verdict drives
- Evidence Reports — obligation evidence per framework, sealed and bundled; the Assurance Report keeps its behavioural role

