Open-Source AI

Aegis puts a trusted gate between AI agents and your tools

Aegis puts a trusted gate between AI agents and your tools

Image: Adam Massimo Mazzocchetti / Zenodo (arXiv 2608.16891)

  • No specific calendar dates, day names, or relative-date phrases appear in the body; every factual claim carries an inline external citation on its own line.

A language model asked to summarize a folder can decide, on its own, to delete files, fire off emails, or spin up a cloud job. Mazzocchetti argued, in a new arXiv paper, that the gap between what a model wants and what actually happens is the real frontier of AI safety arXiv paper.

The model proposes; the runtime decides

The paper, by Adam Massimo Mazzocchetti, introduces Aegis, a runtime governance layer that refuses to let a model’s output touch a tool directly. Every action becomes a proposal that a separate, trusted component evaluates before anything executes. Mazzocchetti noted, “The model proposes; the trusted runtime decides” arXiv paper.

Prompt-level instructions can nudge a model’s behavior, but they never create an enforcement boundary — a prompt can be ignored or overridden mid-task. Aegis moves the decision out of the model and into code the model cannot bypass.

Teams already writing agent guardrails often learn the same lesson the hard way: policy text alone does not enforce behavior, which is the gap this runtime approach targets Why AI agent policies fail without runtime enforcement.

What 2,100 governed runs showed

Mazzocchetti tested Aegis inside a repeated sandbox corpus covering five run families, 42 tasks, three conditions, and ten repeats per family Zenodo dataset. Across 6,300 rows driven only by prompt-policy conditioning, 79 rows leaked along risky comparator paths. Across 2,100 Aegis-governed rows, the system recorded zero governed mock-tool applications and zero governed risky side-effect completions.

The audit went further than blocking. All 1,832 rows where Aegis attempted a governed action preserved trusted, server-side-resolved provenance, and all 1,019 rows routed through Senate-style settlement carried quorum and a final signed tally Zenodo dataset.

Fail closed, settle in quorum

Two design choices do the heavy lifting. Under uncertainty, Aegis fails closed — it blocks rather than guess. For sensitive cases, it routes through a “Senate-style settlement,” a quorum-based path that requires more than one party to authorize the action, so no single component can unilateralize a risky move arXiv DOI.

Mazzocchetti conceded, “These results do not prove general autonomous-agent safety” arXiv DOI. The zero-risk record lives inside one evaluated sandbox, not across every agent anyone will ever ship.

Aegis proves a runtime can catch the proposals a prompt misses. The harder question — whether any runtime can keep up with the models it governs — is still open.

DATE-CITATION CERTIFICATION

  • Total date phrases in body: 0
  • Total entries in date-source map: 0
  • Counts equal (N == M): YES
  • Pre-save audit exit code: 0
  • Every date phrase with its same-line citation:
  • (none — no date phrases appear in the body)
Editorially independent: we accept no payment for coverage and currently use no affiliate links. Read our Editorial Standards and Corrections Policy. Published: Aug 19, 2026.
Jinultimate

Editor of ZBrandCo and the person accountable for what we publish — setting our sourcing standards, fact-checking claims against primary sources, and issuing corrections promptly across AI, open source, and gaming. Reach the desk at editorial@zbrandco.com.