Skip to content
Back to Insights
Agentic AIBy KE Engineering Team

Chain of Thought Is Not an Audit Trail

Chain of Thought Is Not an Audit TrailAgentic AI cover for Chain of Thought Is Not an Audit Trailpass / failAGENTIC AIChain of Thought IsNot an Audit Trail// LOG WHAT HAPPENED

When an agent does something surprising, the instinct is to read its reasoning trace. It said why it did it, right there. The problem is that a model's stated reasoning is a generated artifact, produced by the same process that produced the action, with no guaranteed relationship to the computation that drove the decision.

Three papers show that models produce plausible reasoning that leaves out the factor that decided the output, or that describes a process different from what the activations suggest happened: Language Models Don't Always Say What They Think (2023), Measuring Faithfulness in Chain-of-Thought Reasoning (2023), and Reasoning Models Don't Always Say What They Think (Anthropic, 2025). The model isn't lying. "Explain your reasoning" is just another generation task, and the model is good at generation.

If your audit story is "we log the chain of thought," your audit story is a well-written guess.

What an audit trail needs to answer

After an incident, the questions are concrete. What inputs did the system have. What tools did it call, with what arguments, in what order. What did those tools return. What was in the context window at the moment of each decision. Which policy checks ran and what they decided. Who approved what.

None of those questions are answered by the model's narration. All of them are answered by instrumentation of the system around the model.

Record the observable

typescript
type DecisionRecord = {  step: number;  contextHash: string;          // hash of the exact window at this step  contextSnapshot?: string;     // full window, retained per policy  toolCall?: {    name: string;    args: unknown;              // exactly what was passed    result: unknown;            // exactly what came back    latencyMs: number;  };  policyChecks: { name: string; outcome: "allow" | "deny" | "escalate" }[];  approval?: { by: string; at: string; unchanged: boolean };  // true if the action ran exactly as approved  modelSnapshot: string;        // the exact model version  promptVersion: string;  reasoningText?: string;       // retained, labeled as unverified narration};

The reasoning text is kept. It's useful as a hint. It's labeled for what it is, and nothing downstream treats it as ground truth.

Reproducibility is the test

An audit trail is only as good as your ability to replay from it. Given the context snapshot, the model version, and the prompt version, can you re-run the step and get the same tool call? If the model is non-deterministic at the sampling level, can you at least get the same distribution and confirm the observed action was within it?

If you can't replay, you can't audit. You can only narrate.

Attribution through structure

If you need to know why the model chose tool A over tool B, the reliable approach is counterfactual: re-run with the suspected input removed and see whether the choice changes. It's slow, and it's the most reliable way most teams have to get evidence instead of a story.

For high-stakes actions, you can build attribution in structurally. Require the model to cite which retrieved document or which input field a decision depends on, and verify the citation resolves to something that was in context. A citation to nothing is a hallucination you can detect mechanically.

When narrated reasoning is still worth reading

It's a good debugging hint and a decent prompt-quality signal. Sometimes it shows the model weighing something it shouldn't. It's still not evidence.

An auditor who asks "why did it do that" should get a replayable record, not a paragraph.