Skip to content

Judged evidence & human agreement

Some obligations cannot be recomputed from the log: was the approval rationale consistent with the request? was the notice clear? did the run stay within its intended use? For those, a judged statement runs a tier-B continuous check — a model template over a deterministic, disclosed sample of runs, with an evaluator fingerprint and a budget.

A model’s opinion is not evidence on its own. Three additions make it checkable.

Per check and period, the rubric the judge applies — text, labels, output type, threshold, spec hash — is stored (assurance_check_rubrics) and sealed as an ai.brutor.rubric record whose payload digest covers it. The rubric body never enters a record. Every verdict in the period references the rubric digest. Freezing is idempotent; a changed rubric closes the open period and opens a new one.

2 · Every judged verdict is an assessment record

Section titled “2 · Every judged verdict is an assessment record”

Each judged verdict is an ai.brutor.assessment record (verdict assessment):

  • the value in controls[] — id ai.brutor.judge.<template>, result pass / fail / n/a (met / not met / not applicable);
  • context.check_id and context.evaluator_fingerprint;
  • the rubric digest in references[];
  • chained assesses → the run’s last action record when one exists.

Sealing is best-effort: a seal failure is logged and the check run continues — the verdict is stored, just unsealed, and the statement counts only sealed assessments.

3 · A weekly blind human sample, sealed as adjudications

Section titled “3 · A weekly blind human sample, sealed as adjudications”
  1. Every week a deterministic, stratified, blind sample of judged verdicts is queued for a human (assurance_check_reviews). The sample is seeded by (check, week) — the same inputs always draw the same sample — targets 10 reviews, and is stratified by judged result (at least 2 per stratum where available), so a judge that says “met” 98 % of the time still has its “not met” verdicts reviewed.

  2. While a review is pending, the judge’s answer is hidden from the reviewer (judged_result is withheld in the API response).

  3. The reviewer grades independently: met, not_met or n/a, optionally a score in [0, 1].

  4. The grade seals an ai.brutor.adjudication record: disposition accept/reject (whether the human’s grade agrees with the judge), approver human, human_disposed: true, authority = the review id, chained adjudicates → the assessment. The reviewer’s identity does not enter the record.

Agreement is computed per (check, evaluator fingerprint) over the graded blind reviews of the last 30 days. A fingerprint change resets the sample — a different judge is a different instrument.

Statistic Meaning
Observed agreement share of pairs where judge and human agree
Cohen’s κ agreement beyond chance for categorical results (met / not met / n/a); undefined when both raters used only one category
MAE mean absolute error for numeric scores
Wilson 95 % interval on the observed agreement — honest at small n and at 0 or 1, unlike the normal approximation
insufficient_sample fewer than 10 pairs: the numbers are still shown, but no calibration verdict is drawn

It is stated beside every judged row, never merged into recomputed counts:

JUDGED · met 57 of 60 · human agreement 0.92 [0.83–0.97], n = 34, blind
Inbox item Raised when
judge_uncalibrated observed agreement below 0.80 with a sufficient sample — review debt on the check
calibration_stalled reviews are queued but none has been graded for 14 days

The scheduler freezes rubrics, queues last week’s sample and evaluates calibration hourly.

Template Cited by Statement
no_manipulation EU AI Act Art 5 Sampled runs show no manipulative or deceptive technique
notice_clear EU AI Act Art 50(1) Sampled notices are clear and distinguishable
per_instructions EU AI Act Art 26(1), ISO/IEC 42001 A.9.3/A.9.4 Sampled runs stay within the declared intended use
approval_rationale EU AI Act Art 14 Sampled approval decisions carry a rationale consistent with the request

A check implements judge.<template> when its evaluator’s template field names it. Without a bound check, the statement is not_evaluable.

Compliance → Human Oversight shows the blind review queue and agreement per check; the reviewer grades from there. The API (Control Plane, assurance-check:* permissions):

Method Path Purpose
GET /v1/admin/assurance-checks/{check_id}/reviews {items[ReviewOut], total, counts, agreement}; judged_result withheld while pending
POST /v1/admin/assurance-checks/{check_id}/reviews/sample queue this week’s sample now → {check_id, week, created, items}
POST /v1/admin/assurance-checks/{check_id}/reviews/{review_id} grade: {"human_result": "met" | "not_met" | "n/a", "human_score"?: 0.0–1.0} → {review, agrees_with_judge, seal, agreement}
GET /v1/admin/assurance-checks/{check_id}/agreement {check_id, evaluator_fingerprint, n, agreements, observed_agreement, kappa, ci95, mae, numeric_n, blind, insufficient_sample, below_threshold, threshold: 0.8, window_days: 30, label}

See the API reference for the full shapes.