Skip to content

AI Gateway

At the centre of the platform sits the Brutor AI Gateway: a single entry point that your AI systems, MCP/IDE clients and the User Portal all call, and that fronts every AI resource you allow — models, MCP servers, agent skills, knowledge bases and peer agents.

Brutor AI Gateway architecture: your AI systems, MCP/IDE clients and the Brutor User Portal all enter one Rust proxy. Inside it, the top half is enforced in the request path on every call — identity and RBAC, guardrails, policies, routing and resilience, limits and budgets, and the two-layer cache — and the bottom half is the evidence every call leaves behind: the run ledger, the signed call chain, and drift/replay/liveness signals feeding the assurance report. Below the gateway sit the AI resources it fronts: models, MCP servers, agent skills, knowledgebases and agent peers.

Read the picture in two halves. The enforcement half runs in the request path on every call — it can block, redact or refuse, and there is no opt-out path around it. The evidence half is what every call leaves behind: the run ledger, the signed call chain, and the signals that feed assurance. Governance you can switch off is advice; this is neither optional nor after-the-fact.

Clients don’t change to get this: point an existing OpenAI or Anthropic SDK at the gateway’s base URL and every request is authenticated, policy-checked, metered and logged. The client sees a normal response — governance is invisible until a policy fires.

Fourteen capabilities sit behind that one hop, in four layers. The assurance layer is what the other three feed; the AI layer is what the gateway fronts; the governance layer is what it enforces; the operations layer is what keeps it affordable and visible.

What actually happens to one call, gate by gate:

The path of a single AI call through the gateway. Going out, it passes in order through identity and RBAC, input guardrails, policies, limits and budgets, cache lookup and routing before reaching the provider. Each gate can stop it instead: 401 or 403 with no valid principal, block or redact at guardrails, deny or hold at policies, 429 at a budget or rate cap; a cache hit answers without calling the provider at all. Coming back, the response passes through output guardrails — which can block or redact mid-stream — and is recorded in the run ledger, the usage record and the audit log before reaching the caller. Every outcome, including every refusal, is recorded.

Two properties are worth pausing on:

  • Blocked requests short-circuit. A guardrail violation or exceeded budget stops the call at its gate — the provider is never called, no tokens are spent, and the refusal is recorded with the same fidelity as a success.
  • Every outcome is evidence. Allowed, cached, redacted, denied, held for approval — each lands in the run ledger, the usage record and the audit log. This is what makes assurance possible without instrumenting your application.

For the deployment view behind this picture — the Rust data plane, the Python config plane, the shared database — see Architecture.

Most gateways stop at the call. Assurance is the layer that answers the question a per-request dashboard cannot: is this AI System still doing its job, the way it was approved? It is built entirely from what enforcement already recorded, so nothing has to be instrumented in your application.

The assurance lifecycle. Every governed call is grouped by task into the run ledger with an honest terminal state; baselines learn each system’s normal, and a version-hashed contract records what was declared at sign-off. Those feed six continuously scored signals — liveness, behaviour, reliability, cost, conformance and oversight (whether human approval of gated actions was actually effective) — which combine worst-of into one health verdict: healthy, degraded, drifting, silent or unknown. Too little evidence reads unknown, never green. The verdict drives the fleet board, findings and response, and the assurance report, which ends by stating what it does not cover. Recorded runs are also reusable: replay rehearses a policy change against them with nothing dispatched upstream.

The calls that make up one task, across models, tools, skills and delegated agents, are grouped into a single run with an honest terminal state. Open one and the actions appear in order, with what each cost and how long it took.

The Runs tab of an AI System in the Brutor Admin Console. A table lists runs with their start and finish times, actor, outcome, turns, actions, tools, cost, who closed them and the chain state: two runs abandoned and closed by idle timeout with intact chains, one errored run closed explicitly whose chain reads client asserted, and one completed run costing $0.0583. A completed run is expanded to show the models used in it, gpt-5.2 with four calls and gpt-5.5 with two, each with tokens and cost, followed by the seven actions in order: five model calls and one MCP tool call named doc_get, each with its duration.

Two columns there matter more than they look. Closed by separates a run the client finished from one an idle timeout closed, and chain distinguishes a gateway-signed execution chain from one the client merely asserted, which is the difference between witnessed and reported evidence.

Each signal is scored against the system’s own baseline rather than a threshold you guessed, and each shows its current figure, its distribution and its shape over time, so “cost went up” resolves into a step change, a slow drift or one bad afternoon.

Outcomes separate finished work from work that merely stopped:

The Outcomes signal for an AI System, marked observed. Four counters read completed 2213 at 87.5 percent, completed but degraded 122 at 4.8 percent, errored 114 at 4.5 percent and abandoned 79 at 3.1 percent. Below, a stacked bar chart of run outcomes over time from early June to the first of September shows mostly green completions with a red errored band that grows sharply in the last two days.

Cost per completed run prices the work that actually produced something, and calls out the spend that produced nothing at all:

The cost per completed run signal. Per completed run reads $0.0189 across 2335 completed runs, the median $0.0162 and the 95th percentile $0.0446. A separate figure, spent but nothing produced, reads $3.4995 in orange and is described as errored, abandoned or exhausted runs. A line chart below plots cost per completed run over three months, oscillating around two cents.

Trajectory length shows how many actions a task takes, and whether any run is hitting its ceiling:

The trajectory length signal. Median actions 6, 95th percentile 9, maximum 10, and exhausted 0.0 percent across zero runs. A line chart shows average actions per run holding around six over three months, and a histogram groups runs into buckets: 45 runs of one to two actions, 1014 of three to five, and 1469 of six to ten, with nothing above. A note explains that 79 runs were closed by idle timeout, usually a client that stopped calling rather than a fault in the system.

Tool errors track the half of an agent’s work that is not the model:

The tool errors signal. Tool calls 4569, errors 4, error rate 0.1 percent. A line chart of tool error rate over time sits flat at zero for three months and then rises sharply to nearly 3 percent in the final days.

Liveness is the only signal that fires when traffic stops. Rather than asking you to invent an expectation, Brutor can propose a cadence from the runs it has already seen:

The liveness card on the Signals tab, marked expectation and editable, with a green Alive verdict. Mode is set to continuous, described as expecting a steady stream of runs and alerting if it stops or collapses. The window is one hour, the grace period five minutes and the minimum one run. The last run is timestamped and the system has been quiet for 16 minutes. A button offers to suggest a cadence from observed runs.

Signals describe behaviour in general. Checks are your own conditions, written once, versioned, approved by an operator and then evaluated over every run. Deterministic ones run in the gateway core as a run finalizes:

The continuous checks panel showing six enabled of seven defined. A tier A check named costly run with no approval step is expanded to reveal its expression, run.total_cost_usd greater than 0.10 and approval request count equal to zero, with buttons to backtest seven days or delete it. Below, a list of recent results shows one line per run id with a false verdict and a timestamp.

Checks that need judgement rather than arithmetic are written in plain language and evaluated by a governed judge model on a declared sample of runs:

The same continuous checks panel with a tier B check expanded, named extraction and personal identifiers leaked. Its prompt instructs the judge to answer true if an extracted output contains a personal identity number, full residential address or bank account number belonging to the policyholder. A line beneath states the sampling rate, 25 percent of runs, declared and disclosed, and names the judge model. Recent results list run ids with true or false verdicts and the tokens each judgement cost.

The sampling rate and the judge are stated on the check itself, because a check that inspects a quarter of runs should never be read as covering all of them.

Findings from every source land in one Assurance Inbox with a severity and a state, so detection becomes a queue somebody works rather than a dashboard nobody opens. An open critical finding older than the review SLA is itself a health signal.

The Assurance Inbox showing 8 critical findings, 27 open and 2 acknowledged, filtered to open and acknowledged warnings and criticals. Rows mix categories: a check breach on customer-facing guardrail blocks repeated 102 times, a critical check rate breach on claims triage error rate, several liveness findings naming AI Systems that have gone silent or are stalling, a run failure, and a critical behavioural drift finding on terminal state mix repeated 13 times. Each row carries its severity, category, count, state and timestamp.

From there the response is graduated through the system’s autonomy level rather than a single kill switch: tighten a limit, require approval, restrict, or suspend while a human investigates. A single misbehaving run can be stopped on its own.

A contract is generated from resolved live configuration and pinned by hash, never authored by hand, because a policy document that can disagree with reality is worse than none. Each version expands to show exactly what it froze:

An expanded contract version in the console, v5, pinned by hash, marked active, minted and approved by an operator with the promotion timestamp. Sections list the resources it may use, two models, two MCP servers and one agent identity; the guardrails in force, both marked inherited; what its agents may do, as per-action grants where five actions are allowed and one, updating the claims queue, requires approval; and limits and governance merged down the group chain, including daily and monthly budgets, token and request caps, and temperature and context-window clamps. Autonomy is autonomous, inherited through three groups. A closing note says the version grants nothing new and lists the four caps it tightened.

Per-run ceilings from the contract are checked at the gateway before every call; window bounds such as average cost per completed task are graded in the report and never refuse a call. Every finished run records the contract it ran under, so “what governed this?” has a provable answer months later. Before a change ships, replay rehearses it against recorded traffic offline, dispatching nothing upstream.

The report is the artifact someone outside the platform reads. It gives a verdict in words, the evidence behind it, and an explicit account of its own limits (health and the report).

An Assurance Report in the console with status assured and evidence maturity observable, alongside buttons to capture a snapshot or export JSON or PDF. It states what we know, that the system operated within its approved operating envelope; what prevents full assurance, that 48 percent of runs lack cryptographically signed execution chains and no independent evidence is attached; the recommended action; and why it matters. Five evidence cards show operational evidence and contract assurance as met, execution evidence as partial, and behavioural stability and independent evidence as not met. A health panel reads still learning at 94 percent and lists the six signals, with liveness not applicable and behaviour advisory because the system is classified as an assistant, which relaxes two of six signals. Contract compliance shows tokens, cost ceiling and model calls per run all met. An evidence maturity ladder runs from observable through baseline established, controlled, evidence-backed and independently assured, marking the current rung and what is blocking the next one.

Three things in that picture are deliberate:

  • Unmeasured is reported as unmeasured, never as a pass. A signal the system’s kind relaxes says so, and names the reason.
  • Learned baselines are shown as learned, with how many runs they rest on, so nobody mistakes a fortnight of data for an established norm.
  • The evidence maturity ladder states the rung the system has reached, the specific thing blocking the next one, and what to do about it. Snapshots are immutable, and the whole report exports as JSON or PDF for whoever asked.

What the gateway fronts. Every resource below is bound to a resource group, and every call through it carries the same identity, limits, policy and audit.

One OpenAI-compatible surface across chat, embeddings, images, audio, transcription and video, plus the native Anthropic Messages API at /v1/messages with token counting, so Claude-native SDKs point at Brutor unchanged (LLM API, reference).

  • A preconfigured catalog across the major providers, plus self-hosted Ollama, vLLM and KServe behind the same endpoint, and import of a model’s configuration from Hugging Face.
  • The catalog carries the price list, per token, image, minute and second, which is what lets budgets, caps and cost-based routing work on real numbers.
  • Central model defaults (temperature, top-p, max tokens) per model, so behaviour is consistent without per-call tweaking.
  • Per-model circuit breakers, timeouts, retries and fallback chains, and tracked provider retirement dates with advance operator alerts.
  • Provider keys never leave the gateway: encrypted at rest, resolved server-side, and never exposed to clients (users and API keys).
  • A batch API submits large asynchronous jobs at provider batch rates over the same governed surface.

A routing group is just a model name to the caller: the group picks the real model, so the fleet changes without touching client code.

How a routing group resolves. One model name enters a governed hop where identity, RBAC, budget and input guardrails are checked. Eligibility filters then decide which models may serve it: access, so every member model must be granted to the caller’s group in its own right; cooldown and health; per-model rate limits; and context window. One of five strategies picks from the survivors — weighted random, least busy, lowest usage, lowest latency or lowest cost. A 429, 5xx or timeout retries down the chain and puts the failed model into cooldown; a prompt that is too long is re-routed to a model that fits. The response leaves through output guardrails with spend recorded against the calling group.

  • Five selection strategies: weighted random, least busy, lowest usage, lowest latency or lowest cost, set per group and changed without a redeploy.
  • Failure is a fallback, not an error. A 429, 5xx or timeout retries down the chain and puts the failed model into cooldown; a prompt too long for the chosen model is re-routed to one that fits.
  • Per-model RPM, TPM and concurrency are part of the selection decision, and health checks exclude unhealthy models until they recover.
  • Routing can never widen access. Every member model must be granted to the calling resource group in its own right, and the log names the model that actually served.

Servers are registered once; every tools/call from any client takes the governed hop (register a server, MCP clients).

A tools/call request settles access and cost first — authenticate, server access, quotas, rate limits, data residency — then five independent policy controls decide what the call may do: capability filter, guardrails, semantic policy, argument policy and human approval. Only then does the OAuth proxy attach a vaulted token, the MCP server run the tool, and output guardrails scan the result. Every call becomes one log row feeding the run ledger.

  • A catalog of ready-to-enable servers with images mirrored into Brutor’s own registry, plus bring your own: any container image, repository or remote MCP URL.
  • Virtual MCP aggregation combines tools from many servers behind one endpoint, with per-group filtering and tool renaming, and the MCP Registry publishes and federates records with their governance posture.
  • An OAuth proxy runs the flow per user and keeps real SaaS tokens server-side; the client only ever holds a short-lived gateway JWT.
  • Capability filtering enables individual tools, resources and prompts per group, and each tool is enabled, approval-required or disabled.
  • Guardrails scan arguments on the way in and results on the way back; argument and semantic policies judge what a call would do.
  • Limits apply at group, server and tool scope, and a server’s region is checked against the tenant’s residency profile before it is reached.

A skill packages instructions, scripts, reference files and templates into one versioned bundle: the procedure, not just a prompt (agent skills).

Progressive disclosure across three governed levels. L1 discover returns only the skills the caller’s groups allow, roughly 50-100 tokens each; a skill the caller may not use never appears in the catalog. L2 load returns the full SKILL.md and a manifest of what the skill may read, run and render, with nothing executing and the load written to the audit trail. L3 act runs read_resource, run_script and render_template one step at a time, each step re-governed with guardrails, argument policy, per-skill quota, approval and audit. The bundle never leaves the platform: scripts run sandboxed with no network, a read-only filesystem and a hard timeout.

  • Progressive disclosure over MCP: discover the catalog, load full instructions only when needed, then act step by step, so a skill costs very little context until used.
  • Works from any MCP client, with no Brutor SDK to adopt.
  • The bundle never leaves the platform. Agents receive instructions, reference content and script output, never the code, and scripts run in a sandbox with no network, a read-only filesystem and a hard timeout.
  • Every step is governed independently: guardrails on skill input and output, argument policies, per-skill quotas and approval gates on each action.
  • Access is by resource group, and published versions are immutable, moving through draft, validated, published, deprecated and archived.

Every agent is a principal rather than a shared key: its own identity, a named human owner, and an optional anchor to your IdP (agent identity).

An agent identity is a principal with its own agent-ULID, a named human owner, an IdP anchor, a time-to-live and one-click revocation. Every agent is default-deny. Grants are written per action type — llm_call, mcp_tool, a2a_call, skill_exec — each scoped by a target glob and carrying an effect of allow, approval-required or deny, plus constraints for time window, max delegation depth, rate and expiry. At call time the decision is allowed, approval required, or denied; dry-run mode records what would have been blocked without blocking it.

  • Default-deny. A new agent can call nothing until a grant says otherwise, and nothing is implied by group membership.
  • Grants are written per action (llm_call, mcp_tool, a2a_call, skill_exec) and scoped by a glob over the target, with allow, deny or approval-required, plus time windows, delegation depth, rate and expiry.
  • Dry-run before you enforce: Brutor records what would have been blocked without blocking it.
  • Native A2A v1.0 inbound and outbound (A2A agents), with Ed25519-signed agent cards published at a well-known URL, HMAC-signed delegation chains, the full task lifecycle with streaming, and per-card rate limits.

The retrieval pipeline is already built: chunking, embedding, hybrid search, permission filtering and citation (knowledge bases).

Write path: content from Confluence, Notion, Drive, Slack, GitHub, Jira, SharePoint and a web crawler, or files uploaded directly, reaches the platform via KB Connector Sync or direct upload; the KB Uploader extracts, chunks, embeds as a dense vector plus BM25 tokens, and writes one point per chunk into Qdrant with its source and permissions. Read path: a question arrives, the Gateway Core embeds it with the same model, runs a hybrid dense-plus-BM25 search fused with RRF, filters to chunks the asker could have opened at source, and injects the top passages so the model answers with citations.

  • Qdrant ships in the deployment and you operate it; Brutor manages what lives inside it, including the per-tenant collection and short-lived scoped tokens.
  • Eight connectors (Confluence, Notion, Google Drive, Slack, GitHub, Jira, SharePoint and a web crawler) sync on a schedule, or upload files directly.
  • Hybrid retrieval: dense vectors and BM25, fused with reciprocal rank fusion, so a reworded question and an exact part number both land.
  • Source permissions still apply. A synced document keeps its origin ACL, and retrieval filters to what the asker could have opened themselves.
  • Collections are bound per resource group, answers carry citations, and every retrieval is on the same audit trail as the call it grounded.

Six checks (PII, secrets, prompt injection, jailbreak, toxic content and your own banned words) run inline on every surface (guardrails).

Guardrails run on every surface the gateway carries: LLM input and output, streaming, embeddings, image, audio, batch, MCP tool calls, Skills and A2A. Six checks are available — PII, secrets, prompt injection, jailbreak, toxic content and banned words — and each can be run by the built-in in-process detector or delegated to Microsoft Presidio, AWS Bedrock, Lakera Guard, OpenAI moderation or your own endpoint. Where several run, the most restrictive action wins. If a provider is slow or down, Brutor retries, trips a circuit breaker, falls back to the built-in detector, and finally fails closed.

  • Bring your own detector, per check: the built-in in-process detector, Microsoft Presidio, AWS Bedrock Guardrails, Lakera Guard, OpenAI moderation, or your own endpoint. Built-in means nothing about the request leaves your deployment to be scanned.
  • Most restrictive action wins when several apply: a provider can tighten a decision, never loosen it.
  • Every surface, not just chat: LLM input and output, streaming, embeddings, image, audio, batch, MCP tool calls, skills and A2A. Streaming is guarded either by buffered release or by cutting the stream mid-token.
  • It fails closed. Retry, then circuit breaker, then the built-in detector, then refuse. Break-glass is explicit, time-limited and recorded.

Two engines, one rollout discipline (policies).

Three policy types, each answering a different question. Argument policies ask what a call would actually do, using deterministic SQL, URL, shell, path and JSON analyzers on the model’s own tool calls, MCP input and skill input, rolling out from warn to deny. Semantic policies ask what a request means, using a plain-English constraint judged by a model, rolling out from shadow to enforce. Agent policies ask whether an agent may act at all, using grants per action and target against a default-deny identity, rolling out from dry-run to enforce. All of it is policy-as-code: export, validate, dry-run, apply, then diff and roll back.

  • Argument policies ask what a call would do. Deterministic SQL, URL, shell, path and JSON analyzers read the arguments before anything executes, on the model’s own tool calls, on MCP input and on skill input. Run at warn, promote to deny.
  • Semantic policies ask what it means. A plain-language constraint judged by a model, run in shadow first so real violations are recorded while nothing is blocked.
  • Agent policies ask whether the agent may act at all (see agent identity and agent-to-agent).
  • Policy-as-code, not a wiki page. Export a resource group’s configuration as versioned YAML, validate it against the schema, dry-run to see what would change, apply it as a bundle, then diff two versions and roll back one.

One tree shaped like your org chart, from organization down to a single agent. Every model, server, skill, guardrail, member, key and limit is scoped to a node in it (resource groups, limits).

A resource group tree runs Organization at $1,000/day, Department at $400, Team at $500, and an AI System with no cap of its own. Resources compose additively, gated so a parent can withhold. Limits compose restrictively: the effective daily cap is $400, set at the department — the team asking for $500 does not raise the $400 it sits under. The same most-restrictive rule covers tenant-wide limits set directly on a model, MCP server or skill.

  • Resources compose additively and each inheritance is gated, so a parent can withhold rather than pass everything down.
  • Limits compose restrictively: the effective value is the most restrictive across the node and every ancestor. A $500 cap under a $400 department is still $400, which is what makes delegation safe.
  • Dollar, token and request caps for LLM and MCP traffic alike, enforced in real time: the next call over the line is refused rather than reconciled later.
  • Run-level ceilings for agents: cost, tokens, model calls, delegation depth and active time per run (the clock pauses while a human approval is pending), so one looping agent cannot consume a team’s day.
  • Tenant-wide controls can be set directly on one model, server or skill and merge into the same calculation, and agents are members like anyone else.

Two separate trails: what the AI did, and who changed what it was allowed to do (compliance, audit row reference).

Two sources feed two trails. AI traffic writes the proxy log: who called, which model or tool answered, the decision, tokens and cost, the guardrail or policy that fired, plus compliance tags for declared frameworks. Admin and operator action writes the change trail: every resource carrying its author, editor and version, plus a dedicated operations log for destructive console actions. The proxy log is tamper-evident — each row hashes onto the previous and the writer Ed25519-signs each flushed batch head into a checkpoint. Both trails roll up into the AI Asset Registry, ISO/IEC 42001 coverage, EU AI Act risk classification and GDPR, SOC 2 and HIPAA views, and export as an evidence pack or to S3, Splunk HEC or a JSONL webhook.

  • Proxy logs: one row per call with who called, which model or tool answered, the decision, tokens and cost, and which rule fired, including refusals.
  • Tamper-evident by construction. Each row hashes onto the previous, and the writer Ed25519-signs the head of each flushed batch into a checkpoint that stores the key that signed it, so rotation does not invalidate history.
  • The change trail carries each resource’s author, editor and version, and destructive operator actions get their own rows.
  • Compliance tagging is opt-in: declare GDPR Article 30, SOC 2, HIPAA, EU AI Act or ISO 42001 and tags are written as calls happen; declare none and the engine never runs.
  • Sealed action records. With an evidence key configured, every governed verdict — refusals included — is also sealed as a signed record into a per-tenant Merkle log whose tree heads independent witnesses countersign, so a party outside your organisation can check it with standard COSE/SCITT tooling. See the compliance evidence programme.
  • The Asset Register inventories every model, tool, skill and agent with a fact sheet and append-only snapshots, and evidence exports to S3, Splunk HEC or a JSONL webhook (logs).

Two caches that save in different ways, both surfaced in Mission Control with hit rate and spend avoided (tokenomics).

The response cache tries two layers before the provider is called: an exact SHA-256 hash lookup in Redis at about 0.1 ms, then a semantic cosine-similarity match in Qdrant at about 5 to 50 ms. Either is a hit and returns straight to the client with no tokens spent, still writing an audit row. On a miss the call goes upstream, where the prompt cache applies: Brutor injects the provider’s own cache markers so repeated calls skip prefill, with a cache read costing about 10 percent of input tokens and a write about 125 percent, so break-even is two reads. The two caches treat tools oppositely: the response cache skips tool-using requests, the prompt cache wants them.

  • Response cache, layer 1: an exact hash lookup in Redis. Layer 2: a semantic match in Qdrant above a similarity threshold you set. A hit calls no provider, spends no tokens, and still writes an audit row.
  • Isolated per tenant and per group, with optional sharing for FAQ-style questions, and time bucketing so a “this month” question does not return last month’s answer.
  • Tool-using requests are skipped by default: the answer depended on what a tool returned at that moment.
  • Prompt cache: Brutor injects each provider’s own cache markers so repeated calls skip prefill. A read costs roughly a tenth of the input tokens and a write roughly a quarter more, so break-even is two reads, which is why it belongs on stable prefixes such as tool definitions and long system prompts.

Attribution by provider, model, team and individual user on one screen, with agent and unattributed spend shown as its own line rather than spread across users (Mission Control).

Seven views over one dataset (traffic, cost, LLM economics, governance, caching, agents and reliability) rather than seven disconnected dashboards (metrics, logs).

  • Broken down by surface (LLM, MCP, A2A, Portal) with first-class MCP and A2A views: per-method counts, duration and success rate, and per-capability figures per agent card.
  • Enforcement outcomes as numbers: guardrail blocks as a count and a share, policy denials, detections, and how many requests were inspected at all, which is the difference between “nothing fired” and “nothing was checked”.
  • Discovery coverage as a fraction: governed assets over discovered assets, with shadow AI named, typed and actionable from the list (shadow AI discovery).
  • Un-approved model spend is surfaced with what it cost, and per-framework compliance coverage says plainly when a framework is not enabled.
  • Bring your own stack. Native OpenTelemetry feeds Prometheus, Grafana, Loki or any OTLP backend you already run, and live reachability for every model and MCP server turns a failing dependency into a red component rather than a mysterious error rate.

Nothing here asks you to write against a Brutor SDK. Point an existing OpenAI or Anthropic client at the gateway’s base URL; MCP and A2A are the real protocols, not a dialect (CLI, coding agents). The posture exports as YAML for Git, the register and audit evidence export in standard formats, and the deployment is a single Rust binary on your own infrastructure (architecture).

Not all AI traffic can route through a gateway, and Brutor does not pretend otherwise. Two mechanisms bring the rest into the same picture — observed, costed and attributed, but never enforced:

  • Traffic Data Import calls provider enterprise APIs — usage and spend exports for the ChatGPT, Claude and Copilot seats you buy — and lands them in the same usage ledger as routed traffic. FinOps sees one number per team and provider, whether the tokens went through the gateway or not.
  • Shadow AI Discovery collects the non-routable and unregistered traffic: discovery agents and adapters send signed events about AI usage found on the network, in repositories or in expense data — agents nobody registered, MCP servers nobody governs.