Behavioural drift
Brutor learns what normal looks like for each system from that system’s own history, then reports when the behaviour moves — with the likely cause attached. There is no universal “correct” number of actions per task; the only honest reference point is what this system used to do.
The statistics are deliberately unfashionable
Section titled “The statistics are deliberately unfashionable”Every measure is robust: median and MAD rather than mean and standard deviation, Jensen–Shannon divergence rather than KL, nearest-rank percentiles. Run metrics are heavy-tailed — one agent that loops 400 times moves a mean enough to hide a real shift in the other 500 runs, and a standard deviation computed over that tail is wide enough to swallow anything.
A change must also be sustained, not a single excursion: a metric has to exceed the robust-z threshold and have a located change point. One bad afternoon is not drift.
Learning is not health
Section titled “Learning is not health”Below 50 baseline-eligible runs a system is learning and emits no drift alerts at
all. The UI says “still learning this system’s normal” rather than showing an empty list
that reads as a clean bill of health.
The cause, not the p-value
Section titled “The cause, not the p-value”The alert Brutor tries not to write is:
action_count up 180%, p=0.001
What ships instead is:
action_count up 180% since 14:00 on the 14th; the provider shipped a new point release of the underlying model at 13:52.
Brutor is the only component that sees both the traffic and the model registry, so it is the only thing that can make that correlation. Candidate causes are checked in order and the first that fits wins:
| Rung | What changed | How Brutor knows |
|---|---|---|
incident_window |
The change point lies inside a declared incident window (or within two hours of its edge) — the operator already said what happened | The window: its scope, cause category and reason; the system’s own window is preferred over a tenant-wide one |
contract_change |
A new contract version was promoted | Contract history |
agent_release |
The agent was redeployed — a new release, or the same release behaving like different software | The run’s release and implementation fingerprint |
model_change |
The model behind the system rolled over | Model registry and the run’s models_used |
prompt_change |
The system’s instructions were edited | The dominant instruction fingerprint of runs before and after the change point differs |
dependency_change |
A tool or MCP server was replaced or reconfigured | Server and binding history |
dependency_degraded |
A dependency started failing | Tool-error rate on that tool |
composition_change |
Tools or peers were added or removed from the system | Contract closure |
corpus_change |
The knowledge base it retrieves from changed | The run’s rag_collections set moved |
input_drift |
The inputs themselves changed shape | Prompt-token distribution |
traffic_mix |
A different mix of callers or clients | Client and actor mix |
unexplained |
None of the above lines up | — |
The ladder terminates at unexplained, which is reported as such. An unexplained
finding offers the correlated metrics as a starting point instead of inventing a reason.
prompt_change and corpus_change exist because those are the two changes that ship with
no release at all — the ledger keeps a hash of the instructions and the collection set on
every run precisely so they can be named.
Severity is consequence, not confidence
Section titled “Severity is consequence, not confidence”A statistically bulletproof 3% cost rise is low. A noisy 4× jump in max-iteration
exhaustion on a refund system is critical. Severity comes from effect size × run volume
× business impact (cost delta, autonomy level, EU AI Act risk tier). Ranking by confidence
rather than consequence is how a drift feed becomes noise.
Triage
Section titled “Triage”
Each finding can be acknowledged — I am looking at this; it stays open — or closed. Closing always says which fact it records, and a note: the platform decides nothing silently, so a close without a kind is refused.
Closing a finding is a fact the detector honours
Section titled “Closing a finding is a fact the detector honours”A resolution that only changed a status would contradict itself: the record would say handled while the detector, still reading the same samples, raised the same finding on its next sweep — and a response policy would undo a human restore with no new evidence. So each kind has one defined effect on every signal that reads the affected samples:
| Kind | Means | Effect |
|---|---|---|
accepted_change |
The shift is intended — a release, a retune, a new caller. The new behaviour is the new normal | A baseline epoch starts at the finding’s change point (its detection time when there is none). Samples before it no longer describe normal: the baseline for that metric re-learns from the epoch and reads learning until it has enough new runs — never “ready” on old data. The same shift cannot fire again |
cause_fixed |
Something was wrong and has been fixed — an outage, a bad deploy rolled back | An incident window from the change point to now is declared (bounds, cause category and scope are editable before you confirm). The same samples cannot fire again; new bad behaviour after the fix fires normally |
false_positive |
The detector was wrong | Recorded as not-drift (it feeds detector tuning), and the evidence it misread is not read again for that metric: an epoch starts at the finding’s detection time |
auto_cleared |
Brutor closed it because the condition cleared — an envelope term back within bound, a dependency healthy again, an agent check that no longer finds it | None. Only the platform uses this kind |
Every human resolution is a sealed declaration in the tenant’s evidence log, attributed to the system, with the effect it wrote.
Baseline epochs
Section titled “Baseline epochs”An epoch is keyed like a baseline: system, metric, basis (runs or segments) and caller.
A finding on a service’s aggregate baseline writes caller all, which covers every caller
of that metric. Baselines show the epoch in force — re-learning since … — accepted change
by ….
Measurement changes
Section titled “Measurement changes”Sometimes the platform, not the agent, changes: an upgrade alters how a metric is
measured — for example, a run’s client label changed from the raw User-Agent
(python-httpx/0.28.1) to its product name (python-httpx). Every system’s samples shift on
upgrade day, and without help the detector would report it as the agent changing.
Each such change is declared in the platform’s measurement-change registry. The first time a
release that carries one starts, Brutor writes a baseline epoch of kind measurement_change
(caller all) for every AI System that has a baseline on the affected metric, starting at
that moment, and records that it applied so it never applies twice. The baseline re-learns on
the new measurement, so the shift cannot fire as drift. Baselines show re-learning since … —
measurement changed by the platform: … with the reason, and the assurance report lists every
measurement change under judgement_exclusions.measurement_changes (report spec 1.9).
Incident windows
Section titled “Incident windows”An incident window is a declared period whose samples are excluded from judgement but
kept as evidence. It covers one AI System, or every system in the tenant (a provider
outage), optionally narrowed to a model, provider, MCP server, skill, knowledge base or
agent for display. It carries a reason and a cause category — provider_outage,
platform_incident, maintenance or other. Windows are declared by a cause_fixed
resolution or directly — a planned maintenance — and listed on the AI System and
tenant-wide. Declaring, editing and withdrawing a window are each sealed. Several findings of one
incident close against one window: the first cause_fixed declares it, the rest attach
to it by id (the window must cover their system); each attachment is sealed, and the window
lists every finding it closed — so the window list and the report show each window once.
Windows excuse performance, never facts. Samples in a window are excluded from rate and ratio judgements only: baselines and drift observation windows, envelope terms (rates and per-run ceilings), the health signals’ rate components, the down rule over the latest calls, service signals, connection-line rates, an external dependency’s boundary error rate and latency, and oversight — the share of approval holds decided by a human, on the health signal and Compliance → Human Oversight. Declared-vs-observed facts still see every sample: an undeclared caller or client, a model used outside the contract, an undeclared dependency, an unaccepted caller or a grant without an edge seen only during an outage is still a finding — and a window never excuses an agent or conformance finding. One rule, applied in one place per plane:
- a sample is excluded when its timestamp is inside
[starts_at, ends_at)of a window for its system, or a tenant-wide one — a run by its start; a run segment (a call into a shared service) by its first action, covered by a window on the service or on the run’s root; a gateway call — including an A2A call to an external agent — through its run, so a run is excluded whole, never half; - for a metric with a key, a sample before the latest epoch of its (system, metric, basis)
on its caller or
allis excluded; - an approval hold by when it was raised, against windows for the nearest AI System above its requester’s group, or tenant-wide ones.
A window only excuses what it covers: an approval hold raised outside every window that lapsed because nobody decided it in time (overnight, say) always counts against oversight — that is an approval-window or staffing question, not an incident. No other rule excuses a lapsed hold.
Never hidden: runs inside a window are still recorded, sealed and exportable. The
Assurance Report lists every window — who
declared it, why, how many runs it excluded — and how many approval holds were raised inside
the windows (judgement_exclusions.excluded_approval_requests, report spec 1.10); its
coverage states the excluded period: unmeasured is not 100 %. The approvals ledger still
lists every hold.
Response policies respect the resolution
Section titled “Response policies respect the resolution”A finding closed with any kind stops being a trigger immediately. The response evaluator
never fires on a finding whose evidence — its change point, else its detection time — lies
inside an incident window, or at or before a baseline epoch on the same metric; it records
the refusal (evidence_excluded) instead. A restore after resolution sticks: no
re-detection can come from the same samples, while new evidence after the resolution
still triggers normally.
Systems whose behaviour is legitimately erratic by design can be excluded entirely with
baseline_opt_out — better an explicit switch than an operator muting the whole alert
category.
Related
Section titled “Related”- The run ledger — where the baselines come from (and why partial runs are excluded)
- Autonomy & response — acting on a drift finding automatically
- The gate & replay — drift answers “did it change?” after the fact; replay answers “would this change it?” before

