Judged evidence & human agreement
Some obligations cannot be recomputed from the log: was the approval rationale consistent with the request? was the notice clear? did the run stay within its intended use? For those, a judged statement runs a tier-B continuous check — a model template over a deterministic, disclosed sample of runs, with an evaluator fingerprint and a budget.
A model’s opinion is not evidence on its own. Three additions make it checkable.
1 · The rubric is frozen and sealed
Section titled “1 · The rubric is frozen and sealed”Per check and period, the rubric the judge applies — text, labels, output type, threshold,
spec hash — is stored (assurance_check_rubrics) and sealed as an ai.brutor.rubric
record whose payload digest covers it. The rubric body never enters a record. Every
verdict in the period references the rubric digest. Freezing is idempotent; a changed
rubric closes the open period and opens a new one.
2 · Every judged verdict is an assessment record
Section titled “2 · Every judged verdict is an assessment record”Each judged verdict is an ai.brutor.assessment record (verdict assessment):
- the value in
controls[]— idai.brutor.judge.<template>, resultpass/fail/n/a(met / not met / not applicable); context.check_idandcontext.evaluator_fingerprint;- the rubric digest in
references[]; - chained
assesses→ the run’s last action record when one exists.
Sealing is best-effort: a seal failure is logged and the check run continues — the verdict is stored, just unsealed, and the statement counts only sealed assessments.
3 · A weekly blind human sample, sealed as adjudications
Section titled “3 · A weekly blind human sample, sealed as adjudications”-
Every week a deterministic, stratified, blind sample of judged verdicts is queued for a human (
assurance_check_reviews). The sample is seeded by(check, week)— the same inputs always draw the same sample — targets 10 reviews, and is stratified by judged result (at least 2 per stratum where available), so a judge that says “met” 98 % of the time still has its “not met” verdicts reviewed. -
While a review is pending, the judge’s answer is hidden from the reviewer (
judged_resultis withheld in the API response). -
The reviewer grades independently:
met,not_metorn/a, optionally a score in [0, 1]. -
The grade seals an
ai.brutor.adjudicationrecord: disposition accept/reject (whether the human’s grade agrees with the judge), approverhuman,human_disposed: true, authority = the review id, chainedadjudicates→ the assessment. The reviewer’s identity does not enter the record.
Agreement
Section titled “Agreement”Agreement is computed per (check, evaluator fingerprint) over the graded blind reviews of the last 30 days. A fingerprint change resets the sample — a different judge is a different instrument.
| Statistic | Meaning |
|---|---|
| Observed agreement | share of pairs where judge and human agree |
| Cohen’s κ | agreement beyond chance for categorical results (met / not met / n/a); undefined when both raters used only one category |
| MAE | mean absolute error for numeric scores |
| Wilson 95 % interval | on the observed agreement — honest at small n and at 0 or 1, unlike the normal approximation |
insufficient_sample |
fewer than 10 pairs: the numbers are still shown, but no calibration verdict is drawn |
It is stated beside every judged row, never merged into recomputed counts:
JUDGED · met 57 of 60 · human agreement 0.92 [0.83–0.97], n = 34, blindThresholds and inbox items
Section titled “Thresholds and inbox items”| Inbox item | Raised when |
|---|---|
judge_uncalibrated |
observed agreement below 0.80 with a sufficient sample — review debt on the check |
calibration_stalled |
reviews are queued but none has been graded for 14 days |
The scheduler freezes rubrics, queues last week’s sample and evaluates calibration hourly.
Templates
Section titled “Templates”| Template | Cited by | Statement |
|---|---|---|
no_manipulation |
EU AI Act Art 5 | Sampled runs show no manipulative or deceptive technique |
notice_clear |
EU AI Act Art 50(1) | Sampled notices are clear and distinguishable |
per_instructions |
EU AI Act Art 26(1), ISO/IEC 42001 A.9.3/A.9.4 | Sampled runs stay within the declared intended use |
approval_rationale |
EU AI Act Art 14 | Sampled approval decisions carry a rationale consistent with the request |
A check implements judge.<template> when its evaluator’s template field names it.
Without a bound check, the statement is not_evaluable.
In the console and the API
Section titled “In the console and the API”Compliance → Human Oversight shows the blind review queue and agreement per check; the
reviewer grades from there. The API (Control Plane, assurance-check:* permissions):
| Method | Path | Purpose |
|---|---|---|
| GET | /v1/admin/assurance-checks/{check_id}/reviews |
{items[ReviewOut], total, counts, agreement}; judged_result withheld while pending |
| POST | /v1/admin/assurance-checks/{check_id}/reviews/sample |
queue this week’s sample now → {check_id, week, created, items} |
| POST | /v1/admin/assurance-checks/{check_id}/reviews/{review_id} |
grade: {"human_result": "met" | "not_met" | "n/a", "human_score"?: 0.0–1.0} → {review, agrees_with_judge, seal, agreement} |
| GET | /v1/admin/assurance-checks/{check_id}/agreement |
{check_id, evaluator_fingerprint, n, agreements, observed_agreement, kappa, ci95, mae, numeric_n, blind, insufficient_sample, below_threshold, threshold: 0.8, window_days: 30, label} |
See the API reference for the full shapes.

