← Writing

What Makes an AI Agent Auditable?

TL;DR — Once an LLM agent can call tools and trigger side effects, preventing harm is not enough — its actions must stay answerable after deployment. Auditable Agents argues that no agent system can be accountable without auditability, defines five dimensions of it (action recoverability, lifecycle coverage, policy checkability, responsibility attribution, evidence integrity), and shows why detect, enforce, and recover mechanisms are all needed. Pre-execution mediation with tamper-evident records costs only 8.3 ms median overhead.

LLM agents call tools, query databases, delegate tasks, and trigger external side effects. Most safety work asks whether a harmful action can be prevented. This paper asks a second question that matters as soon as agents are deployed: when something happens, can anyone determine what the agent did, whether it complied, and who is responsible?

Accountability, auditability, auditing

The paper separates three terms that are often blurred:

No agent system can be accountable without auditability.

The five dimensions of agent auditability

  1. Action recoverability — can you reconstruct what the agent actually did?
  2. Lifecycle coverage — does the evidence span pre-deployment, runtime, and post-deployment?
  3. Policy checkability — can an action be checked against an explicit policy?
  4. Responsibility attribution — can an outcome be traced to the agent, tool, or human that caused it?
  5. Evidence integrity — is the record tamper-evident?

Why no single mechanism suffices

Mechanisms fall into three classes — detect, enforce, and recover — and each is limited by when it has information and when it can still intervene. Enforcement acts before execution but sees little context; detection sees more but may be too late to stop the action; recovery works after the fact from whatever evidence survived. In practice a deployed system needs all three.

The evidence

The paper closes with an Auditability Card for documenting agent systems and six open research problems organized by mechanism class. The runtime side of this position is what AEGIS, a pre-execution audit layer, implements.

Frequently asked questions

What is agent auditability?

Auditability is the system property that makes accountability possible: the ability to reconstruct what an AI agent did from trustworthy evidence, check it against policy, and attribute responsibility.

What are the five dimensions of AI agent auditability?

Action recoverability, lifecycle coverage, policy checkability, responsibility attribution, and evidence integrity.

How expensive is runtime auditing for AI agents?

In the paper's runtime feasibility results, pre-execution mediation with tamper-evident records adds only 8.3 ms median overhead.

What is an Auditability Card?

A proposed documentation artifact for agent systems that records how the system addresses each auditability dimension, analogous to model cards for models.


Based on Auditable Agents (KnowFM @ ACL 2026 · ACM AI Summit 2026, arXiv:2604.05485) by Yi Nian, Aojie Yuan, Haiyue Zhang, Jiate Li, Li Li, Xiyang Hu, Hua Wei, Xiongye Xiao, Chaowei Xiao, Yue Zhao. Written by Aojie (Justin) Yuan, USC Fortis Lab.