Skip to content

Behavioural drift

Brutor learns what normal looks like for each system from that system’s own history, then reports when the behaviour moves — with the likely cause attached. There is no universal “correct” number of actions per task; the only honest reference point is what this system used to do.

The statistics are deliberately unfashionable

Section titled “The statistics are deliberately unfashionable”

Every measure is robust: median and MAD rather than mean and standard deviation, Jensen–Shannon divergence rather than KL, nearest-rank percentiles. Run metrics are heavy-tailed — one agent that loops 400 times moves a mean enough to hide a real shift in the other 500 runs, and a standard deviation computed over that tail is wide enough to swallow anything.

A change must also be sustained, not a single excursion: a metric has to exceed the robust-z threshold and have a located change point. One bad afternoon is not drift.

Below 50 baseline-eligible runs a system is learning and emits no drift alerts at all. The UI says “still learning this system’s normal” rather than showing an empty list that reads as a clean bill of health.

The alert Brutor tries not to write is:

action_count up 180%, p=0.001

What ships instead is:

action_count up 180% since 14:00 on the 14th; the provider shipped a new point release of the underlying model at 13:52.

Brutor is the only component that sees both the traffic and the model registry, so it is the only thing that can make that correlation. Candidate causes are checked in order and the first that fits wins:

Rung What changed How Brutor knows
incident_window The change point lies inside a declared incident window (or within two hours of its edge) — the operator already said what happened The window: its scope, cause category and reason; the system’s own window is preferred over a tenant-wide one
contract_change A new contract version was promoted Contract history
agent_release The agent was redeployed — a new release, or the same release behaving like different software The run’s release and implementation fingerprint
model_change The model behind the system rolled over Model registry and the run’s models_used
prompt_change The system’s instructions were edited The dominant instruction fingerprint of runs before and after the change point differs
dependency_change A tool or MCP server was replaced or reconfigured Server and binding history
dependency_degraded A dependency started failing Tool-error rate on that tool
composition_change Tools or peers were added or removed from the system Contract closure
corpus_change The knowledge base it retrieves from changed The run’s rag_collections set moved
input_drift The inputs themselves changed shape Prompt-token distribution
traffic_mix A different mix of callers or clients Client and actor mix
unexplained None of the above lines up —

The ladder terminates at unexplained, which is reported as such. An unexplained finding offers the correlated metrics as a starting point instead of inventing a reason. prompt_change and corpus_change exist because those are the two changes that ship with no release at all — the ledger keeps a hash of the instructions and the collection set on every run precisely so they can be named.

A statistically bulletproof 3% cost rise is low. A noisy 4× jump in max-iteration exhaustion on a refund system is critical. Severity comes from effect size × run volume × business impact (cost delta, autonomy level, EU AI Act risk tier). Ranking by confidence rather than consequence is how a drift feed becomes noise.

The Drift tab on a system: a finding — run outcome mix shifted, seen four times, severity low — with its cause line reading no cause identified and the honest note that no change to the model, tools, composition or inputs lines up with it, and the three triage buttons. Beside it, Automatic response shows no policy armed, so nothing happens without a human, and the recent responses record each finding that matched no policy.

Each finding can be acknowledged — I am looking at this; it stays open — or closed. Closing always says which fact it records, and a note: the platform decides nothing silently, so a close without a kind is refused.

Closing a finding is a fact the detector honours

Section titled “Closing a finding is a fact the detector honours”

A resolution that only changed a status would contradict itself: the record would say handled while the detector, still reading the same samples, raised the same finding on its next sweep — and a response policy would undo a human restore with no new evidence. So each kind has one defined effect on every signal that reads the affected samples:

Kind Means Effect
accepted_change The shift is intended — a release, a retune, a new caller. The new behaviour is the new normal A baseline epoch starts at the finding’s change point (its detection time when there is none). Samples before it no longer describe normal: the baseline for that metric re-learns from the epoch and reads learning until it has enough new runs — never “ready” on old data. The same shift cannot fire again
cause_fixed Something was wrong and has been fixed — an outage, a bad deploy rolled back An incident window from the change point to now is declared (bounds, cause category and scope are editable before you confirm). The same samples cannot fire again; new bad behaviour after the fix fires normally
false_positive The detector was wrong Recorded as not-drift (it feeds detector tuning), and the evidence it misread is not read again for that metric: an epoch starts at the finding’s detection time
auto_cleared Brutor closed it because the condition cleared — an envelope term back within bound, a dependency healthy again, an agent check that no longer finds it None. Only the platform uses this kind

Every human resolution is a sealed declaration in the tenant’s evidence log, attributed to the system, with the effect it wrote.

An epoch is keyed like a baseline: system, metric, basis (runs or segments) and caller. A finding on a service’s aggregate baseline writes caller all, which covers every caller of that metric. Baselines show the epoch in force — re-learning since … — accepted change by ….

Sometimes the platform, not the agent, changes: an upgrade alters how a metric is measured — for example, a run’s client label changed from the raw User-Agent (python-httpx/0.28.1) to its product name (python-httpx). Every system’s samples shift on upgrade day, and without help the detector would report it as the agent changing.

Each such change is declared in the platform’s measurement-change registry. The first time a release that carries one starts, Brutor writes a baseline epoch of kind measurement_change (caller all) for every AI System that has a baseline on the affected metric, starting at that moment, and records that it applied so it never applies twice. The baseline re-learns on the new measurement, so the shift cannot fire as drift. Baselines show re-learning since … — measurement changed by the platform: … with the reason, and the assurance report lists every measurement change under judgement_exclusions.measurement_changes (report spec 1.9).

An incident window is a declared period whose samples are excluded from judgement but kept as evidence. It covers one AI System, or every system in the tenant (a provider outage), optionally narrowed to a model, provider, MCP server, skill, knowledge base or agent for display. It carries a reason and a cause category — provider_outage, platform_incident, maintenance or other. Windows are declared by a cause_fixed resolution or directly — a planned maintenance — and listed on the AI System and tenant-wide. Declaring, editing and withdrawing a window are each sealed. Several findings of one incident close against one window: the first cause_fixed declares it, the rest attach to it by id (the window must cover their system); each attachment is sealed, and the window lists every finding it closed — so the window list and the report show each window once.

Windows excuse performance, never facts. Samples in a window are excluded from rate and ratio judgements only: baselines and drift observation windows, envelope terms (rates and per-run ceilings), the health signals’ rate components, the down rule over the latest calls, service signals, connection-line rates, an external dependency’s boundary error rate and latency, and oversight — the share of approval holds decided by a human, on the health signal and Compliance → Human Oversight. Declared-vs-observed facts still see every sample: an undeclared caller or client, a model used outside the contract, an undeclared dependency, an unaccepted caller or a grant without an edge seen only during an outage is still a finding — and a window never excuses an agent or conformance finding. One rule, applied in one place per plane:

  • a sample is excluded when its timestamp is inside [starts_at, ends_at) of a window for its system, or a tenant-wide one — a run by its start; a run segment (a call into a shared service) by its first action, covered by a window on the service or on the run’s root; a gateway call — including an A2A call to an external agent — through its run, so a run is excluded whole, never half;
  • for a metric with a key, a sample before the latest epoch of its (system, metric, basis) on its caller or all is excluded;
  • an approval hold by when it was raised, against windows for the nearest AI System above its requester’s group, or tenant-wide ones.

A window only excuses what it covers: an approval hold raised outside every window that lapsed because nobody decided it in time (overnight, say) always counts against oversight — that is an approval-window or staffing question, not an incident. No other rule excuses a lapsed hold.

Never hidden: runs inside a window are still recorded, sealed and exportable. The Assurance Report lists every window — who declared it, why, how many runs it excluded — and how many approval holds were raised inside the windows (judgement_exclusions.excluded_approval_requests, report spec 1.10); its coverage states the excluded period: unmeasured is not 100 %. The approvals ledger still lists every hold.

A finding closed with any kind stops being a trigger immediately. The response evaluator never fires on a finding whose evidence — its change point, else its detection time — lies inside an incident window, or at or before a baseline epoch on the same metric; it records the refusal (evidence_excluded) instead. A restore after resolution sticks: no re-detection can come from the same samples, while new evidence after the resolution still triggers normally.

Systems whose behaviour is legitimately erratic by design can be excluded entirely with baseline_opt_out — better an explicit switch than an operator muting the whole alert category.

  • The run ledger — where the baselines come from (and why partial runs are excluded)
  • Autonomy & response — acting on a drift finding automatically
  • The gate & replay — drift answers “did it change?” after the fact; replay answers “would this change it?” before