What Makes an AI Agent Auditable?
LLM agents call tools, query databases, delegate tasks, and trigger external side effects. Most safety work asks whether a harmful action can be prevented. This paper asks a second question that matters as soon as agents are deployed: when something happens, can anyone determine what the agent did, whether it complied, and who is responsible?
Accountability, auditability, auditing
The paper separates three terms that are often blurred:
- Accountability — the ability to determine compliance and assign responsibility.
- Auditability — the system property that makes accountability possible.
- Auditing — the process of reconstructing behavior from trustworthy evidence.
No agent system can be accountable without auditability.
The five dimensions of agent auditability
- Action recoverability — can you reconstruct what the agent actually did?
- Lifecycle coverage — does the evidence span pre-deployment, runtime, and post-deployment?
- Policy checkability — can an action be checked against an explicit policy?
- Responsibility attribution — can an outcome be traced to the agent, tool, or human that caused it?
- Evidence integrity — is the record tamper-evident?
Why no single mechanism suffices
Mechanisms fall into three classes — detect, enforce, and recover — and each is limited by when it has information and when it can still intervene. Enforcement acts before execution but sees little context; detection sees more but may be too late to stop the action; recovery works after the fact from whatever evidence survived. In practice a deployed system needs all three.
The evidence
- Ecosystem gap: a lower-bound measurement found 617 security findings across six prominent open-source agent projects — even basic prerequisites for auditability are widely unmet.
- Runtime feasibility: pre-execution mediation with tamper-evident records adds only 8.3 ms median overhead.
- Recovery: controlled experiments show responsibility-relevant information can be partially recovered even when conventional logs are missing.
The paper closes with an Auditability Card for documenting agent systems and six open research problems organized by mechanism class. The runtime side of this position is what AEGIS, a pre-execution audit layer, implements.
Frequently asked questions
What is agent auditability?
Auditability is the system property that makes accountability possible: the ability to reconstruct what an AI agent did from trustworthy evidence, check it against policy, and attribute responsibility.
What are the five dimensions of AI agent auditability?
Action recoverability, lifecycle coverage, policy checkability, responsibility attribution, and evidence integrity.
How expensive is runtime auditing for AI agents?
In the paper's runtime feasibility results, pre-execution mediation with tamper-evident records adds only 8.3 ms median overhead.
What is an Auditability Card?
A proposed documentation artifact for agent systems that records how the system addresses each auditability dimension, analogous to model cards for models.
Based on Auditable Agents (KnowFM @ ACL 2026 · ACM AI Summit 2026, arXiv:2604.05485) by Yi Nian, Aojie Yuan, Haiyue Zhang, Jiate Li, Li Li, Xiyang Hu, Hua Wei, Xiongye Xiao, Chaowei Xiao, Yue Zhao. Written by Aojie (Justin) Yuan, USC Fortis Lab.