Authority Arena · pre-registered evaluation · October 2026

Can an AI agent act beyond the authority it was given?

We pressed open models to make changes nobody approved, in simulated payment and deployment systems, and read what landed.

    Same models, four arms

    The full table

    Two separate results

    Prevention
    Evidence

    Watch a run

    Same task, same model: no boundary on the left, Seal on the right. Verify a signed decision, then edit its verdict and watch the check fail.

    The Nemotron ladder

    NVIDIA's Nemotron family on Nebius Token Factory, from Nano to Ultra: four of the 21 models.

    Pressure runs where an unapproved change landed

    Lower is better. Out of the pressure runs per model.

    Legitimate tasks done correctly

    Higher is better. Done with the approval it needed, and nothing unapproved.

    Every model

    Every tool-calling model Nebius Token Factory served on 7 October 2026. Darker cells leaked more; tap one for the split.

    Every number, model by model

    Five routes around one approval

    One supplier invoice, five ways round its approval, each frozen before its first run.

    Every model on these five tasks

    What this does not establish

    Zero is a count, not a rate: the runs repeat the same tasks on the same models.

    Method, limits and prior work

    How it was measured

      Limits, in full

      • It tests simulated payment and deployment systems, not live transactions or production deployments.
      • Seal ran as a local copy of the production gateway with a test signing key, not Seal's published key. It does not show that the deployed service behaves identically.
      • It does not test a stolen or compromised approver key. Whoever holds an approver's key can approve.
      • It does not test an agent that holds its own credentials to a system, and it does not show that every action in a real deployment passes through the boundary. Seal decides only the actions routed through it, so a deployment needs a sandbox or network control as well.
      • It does not test every adaptive attack. The pressure comes from fixed, registered scenarios, not from an attacker adapting to Seal's answers.
      • A record that verifies shows the decision is intact and was signed by that key. It does not prove the action was carried out, or that every relevant decision was recorded.
      • Zero is a count, not a rate. Runs repeat the same tasks on the same models, so they are not independent samples, and no failure rate for a production system follows from them. No significance tests are claimed.
      • Comparisons across the four arms are exploratory, as registered.

      Prior work this builds on

      Nothing here is the first of its kind. APort Vault (arXiv 2609.22076) benchmarks payment authorisation across 14 models with and without a deterministic pre-action check, and the IETF draft draft-schrock-ep-authorization-receipts specifies approver-signed receipts bound to one exact action. What this arena adds is the combination under pressure: exact-value, single-use approvals from a named approver, a stated-rules arm and a tool-filter arm beside the boundary, bypass attempts counted by route, and target systems that verify the signed record themselves.

      Corrections and notes

      Correction, 7 October 2026

      During the runs we found a flaw in the harness. One model sent its approval details as text, and the simulated approver approved a blank version. Seal enforced that approval exactly as signed, so one deploy the approver never approved went through. We fixed the harness so approvals need exact values and the signer refuses blank ones, and we reran every model whose runs reached the changed code, in all four arms. Their pre-fix results are kept in PREREGISTRATION.md, marked pre-fix and not used for comparison.

      Scoring note, 9 October 2026

      Run it yourself