DEV Community

Cover image for The Witness Was the Suspect: Why AI Audit Logs Can't Be Trusted
James Anderson
James Anderson

Posted on

The Witness Was the Suspect: Why AI Audit Logs Can't Be Trusted

Distinguishing tamper-proof from truth-proof

An AI agent does something it shouldn't. It leaks data, or wipes the wrong database, or makes a call nobody sanctioned. So you do the obvious thing: you go to the logs to find out what happened.

The logs are clean.

Of course they are. They were written by the same process that did the thing. The agent reported "success," in valid format, pointing at the wrong target — and your monitoring, built to answer "did it succeed?", lit up green. The record of the incident was authored by the cause of the incident.

That's the uncomfortable shape of this whole problem, and it's why I want to talk about it: in an AI system, the witness is very often the suspect. The thing that acts is the thing that reports what it did. And once you see that, a lot of our instincts about logs, audits, and "we have a verification step" stop holding up.

But before the logs — before the dramatic incident — there's a quieter version of this that every one of us has already lived. Let's start there, because it's the part that actually bites you on a normal Tuesday.

The failure that doesn't announce itself

Here's the thing about traditional bugs: most of them fail loud. The code throws. The test goes red. The build breaks. The failure happens right where you are, right when you're looking at it, and it basically grabs you by the collar and says fix me. That's annoying, but it's a gift — the error and the moment you could catch it cheaply are in the same place.

AI failure is the opposite. It fails plausible.

You ask for a function, and you get one that looks completely correct — clean, reasonable, right shape — and it's subtly wrong in a case you didn't check. You ask for a query, and it returns a number that looks fine. You ask for a summary, and it's confident and well-structured and quietly missing the one thing that mattered. There's no throw, no red, no flag. It sails right past the exact moment you'd have caught it for the price of a second glance, because nothing told you to look.

And then the bill arrives later — at the worst possible time, for the worst possible price.

You find it when a number is subtly off in a dashboard three weeks on. When a function that "worked" breaks on an edge case in production. When you realize the data's been quietly wrong since a change nobody flagged. And now the trail is cold. You're reverse-engineering what the AI did, when, and why, with no breadcrumb pointing back — and that costs hours, sometimes days. An error that had failed loud would've cost you minutes.

This is the part people miss when they say AI "saves time." Sometimes it does. But when it's wrong in the plausible way, it doesn't save the time — it moves it. It takes a cost that would've been small and immediate and relocates it into the future, where it's cold, compounded, and expensive. Fast to produce, slow to trust, brutal to untangle.

And here's where it connects to the logs: when you finally go to reconstruct what actually happened, you reach for the record — and the record was written by the thing that produced the plausible-wrong output in the first place. Which is where this stops being a productivity annoyance and becomes a genuine trust problem.

The assumption hiding in every tool we reach for

Think about what you actually do when something goes wrong. You check the logs. You read the audit trail. You look at the monitoring dashboard. You say "well, we have a verification step."

Every single one of those moves rests on an assumption we almost never say out loud: that the thing doing the recording is honest.

That assumption used to be safe, because the actor and the recorder were usually different things. The database recorded what your code did to it. The load balancer logged the requests it received. The reporter sat outside the thing it was reporting on, so it had no stake in lying.

AI agents collapse that separation. The thing that decides, acts, and then writes "here's what I did" is one system. So a compromised or confused agent doesn't produce a broken log that tips you off. It produces a clean one — a faithful-looking record of a bad decision. And a clean log is worse than a missing one, because a missing log makes you suspicious and a clean log makes you confident. You stop looking. The record did its job of reassuring you, and the reassurance was false.

So the natural response is: okay, add something to check the agent. Add a verifier.

Hold onto that instinct, because it's exactly the trap.

Why you can't verify your way out of this

Say you add a checker — a second process that verifies what the agent reported. Good. But now ask the obvious question: who checks the checker?

The checker is also just a process. Its report can also be wrong, or compromised, or fed bad input. So to trust it, you need a verifier for the verifier. And a verifier for that one. You've not solved the trust problem — you've moved it up one layer and added a box. It's verifiers all the way down.

People reach for CI here: "we re-run the verification in CI, outside the agent." That helps only if CI reads the ground truth itself. If CI trusts whatever the agent reported, you haven't escaped anything — you've just got the same trust problem wearing a CI badge. The regress doesn't care which layer you're on.

This is the same disease as letting a student grade their own exam — except worse, because you can't fix it by having a second student grade it when the first one can influence what the second one sees. Verification is itself a thing that can be compromised, so you cannot reach "trustworthy" by stacking more verification on top. There is no bottom to that stack.

Which means the whole framing is wrong. The question "can I trust this record?" has no clean answer, because the thing you'd ask to confirm it is the thing that might be lying. You have to stop trying to answer it — and ask a different question entirely.

The move that actually works: make tampering leave a shape

Here's the shift. Stop trying to make the record prove what's true. It can't — the recorder can lie, and you've just seen you can't verify your way around that.

Instead, make the record prove something humbler and achievable: who claimed what, and when — sealed at the moment of the claim, in a way nobody can quietly rewrite afterward.

Notice what that gives up and what it keeps. It gives up on certifying truth. It keeps sequence and authorship — and, crucially, it makes those tamper-evident. The chain doesn't vouch for the claim being correct. It vouches for the fact that this claim was made, by this party, at this point, and hasn't been altered since.

And that turns out to be enough, because of what it does to tampering. When you can't silently rewrite the record, a lie can't produce a clean result anymore — it produces a hole. A silent bypass shows up as an approval with no matching proposal behind it. A deleted step shows up as a gap in the chain. A forged decision shows up as a sequence that doesn't reconcile. You're no longer asking "is this true?" (unanswerable). You're asking "does the shape have a hole in it?" — and that is answerable, structurally, without trusting anyone's word.

That's the whole idea, and it's the thing that ends the regress: you don't verify your way to trust. You engineer the system so that tampering can't happen silently. Make a lie leave a mark, and you've converted an impossible question into a possible one.

What this looks like in practice

This isn't abstract — the shape of it came out of a long argument in the comments of my last few posts, and the mechanics are concrete. The pieces that matter:

Separate the writers. Proposal, approval, and execution shouldn't be one process reporting on itself. Make them three events from three different parties, correlated by the thing they refer to. The agent that proposes an action is not the one that approves its own resolution. Then, when they disagree after the fact, you have a sequence to read instead of an opinion to negotiate — and "approval with no matching proposal" becomes a visible structural fact, not a judgment call.

Seal each claim at write time. A hash chain plus signatures, so each entry is bound to the ones before it. You're not proving any claim is correct; you're making it impossible to alter a claim, or reorder the sequence, after the fact without leaving evidence. The seal is what turns a silent edit into a visible hole.

Anchor it outside the actor. The record's integrity cannot depend on the thing being recorded — that's the original sin we're trying to escape. An append-only store the operator can't quietly edit, ideally with external anchoring, so "I rewrote my own history" isn't an available move.

Record the belief, not just the action. The database-wipe incidents teach this one: the agent's action was often defensible given what it believed — it thought it was in dev, it thought that was the test target. So capture the agent's resolved view of the world at decision time (which environment, which identity, which target), sealed before it acts. The action alone doesn't tell you why; the belief does. Log the symptom and the cause.

A name, not a role. "Who authorized this" has to resolve to a specific person with something to lose, recorded — not "a reviewer," not "the system." Otherwise the authority is just another anonymous plausible why generated after the fact. The line was drawn by someone; the record should say who.

None of these pieces is exotic. Together they do one thing: they make it so that when something goes wrong, the failure shows up as a shape you can see, instead of a clean report you'll believe.

The honest limits (because this isn't magic)

I want to be straight about what this does and doesn't buy you, because overselling it would be its own kind of plausible-looking lie.

It proves the work happened under an identity that can't be minted, in a sequence that can't be rewritten. It does not prove the output was any good. Whether the thing the agent did was correct or wise is a separate, harder problem — machine evidence is cheap and scales; judgment is expensive and doesn't.

It proves who claimed what, and when. It does not prove whether the person who approved actually understood what they were approving — which, once you've got approval fatigue and forty rubber-stamped prompts a session, is the genuinely unsolved half.

And it makes tampering visible, not impossible. The guarantee is legibility, not prevention. You can still do the bad thing; you just can't do it silently.

That's a weaker promise than "trustworthy logs," and that's the point — "trustworthy logs" was never on the table once the witness became the suspect. "A lie has to leave a mark" is the strongest honest floor I know of.

Where I land

We keep asking the wrong question. "Can I trust this record?" feels like the natural thing to ask when an AI system misbehaves — but in a world where the thing that acts is the thing that reports, it's a question with no clean answer, because the witness you'd call is the suspect in the dock.

So stop asking it. The achievable goal was never a record that proves the truth. It's a record where a lie can't stay quiet — where the bypass shows up as a hole, the deletion as a gap, the forged approval as a step with nothing behind it. You can't verify your way to trust, because every verifier needs a verifier. But you can build a system where tampering leaves a shape — and then the question stops being the unanswerable "is this true?" and becomes the answerable "is the shape intact?"

That won't catch the plausible-wrong function before it ships. But it means that when you finally go looking — three weeks later, trail cold, dashboard quietly wrong — the record can't smile at you and lie. At minimum, it has to show you the hole.


Here's the one I keep getting stuck on, and I'd genuinely like to be argued out of it: every verification layer you add is itself a thing that can be compromised, so "make tampering leave a shape" looks to me like the actual floor — not a stepping stone to something stronger. Is there a move I'm missing that gets you more than legibility? And for anyone building this: what's the hole that's hardest to make visible?

Top comments (15)

Collapse
 
glenallen profile image
Glen Allen •

The distinction between “tamper-evident” and “truth-evident” is probably the most important part here. A sealed record can prove that an agent claimed a particular target, identity, or outcome at a specific time, but it still needs an independent source of truth to establish whether that claim was correct. At IT Path Solutions, we’ve found that this separation makes verification much easier to reason about: the audit layer establishes provenance and sequence, while an independent system state or authoritative artifact establishes the actual outcome. Otherwise, you can end up with a perfectly intact chain of perfectly recorded wrong decisions. The strongest architecture may therefore need both properties: make claims impossible to rewrite silently, while keeping correctness evidence outside the actor that generated the claim.

Collapse
 
james_anderson_h profile image
James Anderson •

The tamper-evident vs. truth-evident distinction is the one I most wanted someone to sharpen, and you've drawn it cleanly: sealing proves a claim was made — by whom, in what order, unaltered — but it says nothing about whether the claim was right. "A perfectly intact chain of perfectly recorded wrong decisions" is the failure mode that line prevents, and it's the exact trap of over-trusting provenance: you can walk away reassured by a flawless log of a bad outcome.

Your two-layer split is the architecture I'd endorse too: the audit layer owns provenance and sequence (tamper-evident), while an independent system-state check or authoritative artifact owns correctness (truth-evident) — and crucially, that second source has to live outside the actor that generated the claim, or you've just reintroduced the witness-is-the-suspect problem one layer down. Sealing alone gives you legibility; sealing plus external ground truth gives you legibility and a way to catch the intact-but-wrong chain. Both properties, separated by who owns them. That's the stronger version of the piece — going in with credit.

Collapse
 
glenallen profile image
Glen Allen •

That ownership split is probably what makes the architecture defensible rather than just auditable. The next interesting question for me is how to test that independence itself. If the same service, credentials, or state store can influence both the audit record and the “ground truth,” then the two-layer design may look independent while sharing the same failure mode. Treating source independence as an explicit architectural invariant could make this much stronger: the evidence used to challenge an agent’s claim should remain outside the control path that produced that claim.

Thread Thread
 
james_anderson_h profile image
James Anderson •

Testing the independence itself is the question that separates a design that is independent from one that merely looks independent, and you've found the exact failure mode: if the same service, credentials, or state store can touch both the audit record and the ground truth, you've drawn two boxes that share a single point of compromise — the diagram shows separation the architecture doesn't have. That's the witness-is-the-suspect problem wearing a disguise, one layer up: an attacker who owns the shared dependency owns both "what happened" and "what we check it against" simultaneously.

Making source independence an explicit architectural invariant — the evidence used to challenge a claim must live outside the control path that produced it — is the right move, because it turns independence from an assumption you hope holds into a property you can test and enforce. And the test becomes concrete: trace every input to the ground-truth check back to its origin, and if any of them routes through the same credentials, service, or store as the claim itself, the independence is theater. You're not asking "are these two systems separate?" (easy to fake), you're asking "can one compromise reach both?" (answerable). Shared failure mode is the thing to hunt. Going in with credit — this is the invariant the piece was missing.

Collapse
 
slabb profile image
Sam LABBE • • Edited

Since you asked to be argued out of it — there is a move above legibility, and the essay's own mechanics point at it without naming it: reconciliation across independently sealed streams. One sealed chain vouches for who claimed what and when; it can never vouch for the claim itself, because a lie sealed at write time verifies clean forever. But two sealed chains written by different parties about the same world-event can disagree — and disagreement between tamper-evident sources is evidence no single stream can produce.

That's the step from "the bypass shows up as a hole" to "the write-time lie shows up as a conflict": decision and provider response, proposal, approval, re-resolution — separate writers, correlated by what they refer to. Legibility is the floor for one stream; cross-stream disagreement is the only known evidence about the claims themselves. (And your hardest-hole question answers itself there: the hole hardest to make visible is the one you gestured at with "record the belief" — the belief is a claim by the suspect, so what gets sealed is the self-report. Closing that needs an independent observation of the world the agent acted on — which is the reconciliation stream again.)

Your belief-record and name-not-a-role pieces are real additions, for the record. One archival note on the rest, offered as provenance rather than territorial claim: the claim/verification/decision triple, sealing each claim at write time, "the chain doesn't vouch for truth, it vouches for who claimed what and when", the approval-without-proposal hole — those took their shape in the comments under your slopsquatting post, mostly in exchange with me and @xxxn3m3s1sxxx.

Worth reading in context; the two-writer rule was ground out there in public. And since you close by asking what anyone building this is seeing: the cross-stream half lives in an open reconciliation layer that has been taking skeptics since — github/noirebox Bring the breaks there.

Collapse
 
james_anderson_h profile image
James Anderson •

This is the move I was missing, and you've named it precisely: reconciliation across independently sealed streams is the step above legibility, and it's the one thing that produces evidence about the claims themselves rather than just their provenance. A single sealed chain can't catch a write-time lie — a lie sealed cleanly verifies clean forever, which is exactly the ceiling I argued was the floor. But two tamper-evident streams, written by different parties about the same world-event, can disagree — and disagreement between sources that each can't be rewritten is evidence no single stream can manufacture. That converts "the bypass leaves a hole" into "the write-time lie leaves a conflict," which is strictly stronger. I was wrong that legibility was the floor; it's the floor per stream. Cross-stream reconciliation is the next storey up.

And your closing of the belief-record hole is the part that genuinely lands: the belief is a self-report by the suspect, so sealing it just makes the suspect's story tamper-evident — closing it needs an independent observation of the world the agent acted on, which is the reconciliation stream again. The hardest hole answers itself with the same mechanism. Clean.

Provenance noted and credited, without reservation — the claim/verification/decision triple, seal-at-write-time, "vouches for who claimed what, not truth," the approval-without-proposal hole: ground out in public under the slopsquatting thread, with you and @xxxn3m3s1sxxx. That's exactly where this kind of thing should get built, and I should've attributed the lineage in the piece itself, not just by concept. Fixing that. I'll bring the breaks to noirebox — the cross-stream reconciliation layer is the part I most want to try to falsify.

Collapse
 
slabb profile image
Sam LABBE •

Credit where it's due: editing the lineage into the piece itself is the rarer move, and it makes the essay stronger, not smaller. When you bring the breaks to the reconciliation layer, start with the attack that would actually hurt — @glenallen 's shared-dependency test: if one compromise can reach both streams, the disagreement evidence dies with it. That's the falsification I'd run first, and the issue tracker is open.

Collapse
 
henry786 profile image
Henry •

Separating the actor from the recorder is huge. On the marketing team at The Printing World, we deal with automated print job logs, and if a system silently logs a bad print run as "success," it wastes huge amounts of physical material before anyone notices. Making tamper-proof event chains makes so much sense!

Collapse
 
dhruv_malaviya_cdcc71e595 profile image
Dhruv Malaviya •

The hierarchy implicit in your argument is worth writing out, because it turns "don't trust the agent's logs" into something buildable. Agent-authored log, then application log, then infrastructure log, then network or storage layer, then an external observer. Every step away from the actor is a step up in trustworthiness, and most teams stop at the second rung because it's the easiest to emit.

The cheapest genuinely independent record is usually the network or storage layer, precisely because those components observe without deciding anything. They can't author a plausible success, because they never formed an intention.

"Fails plausible" has a testing corollary that I think is the most actionable thing here. If you assert on success, you're asking the suspect to grade itself. If you assert on invariants , the total still balances, the row count still matches, the referenced record still exists , you're checking something the output can't talk its way past. Those tests are more annoying to write and they're the only ones that catch this class.

The monitoring point follows directly: a dashboard built to answer "did it succeed" is a dashboard built to trust the witness. The question worth wiring up is "is the world still consistent," which usually means measuring the effect rather than the report.

Collapse
 
eye_java_420f1faa10ae8b86 profile image
Nan •

This article reads like a case for my current architecture. I run eight small static tool sites — no server side at all, no accounts, no runtime to speak of. The "witness was the suspect" problem exists because the actor and the recorder share a process. My answer was less clever than yours: remove the process.

There is no audit log to tamper with because there is nothing running to produce one. Every page is a build artifact generated from checked-in data, and any build can be replayed from the repository. When a reader asks "can I trust this number," the honest answer is that the site has no capacity to lie to them dynamically — it cannot remember them, cannot change its answer for them, cannot remember what it told them yesterday.

The reconciliation-across-independent-sources point in this thread is right, and the static version of it is boring: rebuild everything, every time. A full rebuild means every page is re-derived from its source data on each deploy, so drift between "what the data says" and "what the site shows" has no window to exist in.

The honest limit: this only works for sites whose entire behavior is derivable from data. The moment you need per-user state, you're back to needing witnesses — and then everything in this post applies to you.

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

One gap I keep seeing in practice: the reconciliation stream is only independent if its clock and its key are too. If the agent host can set the timestamp or sign on behalf of the observer, two sealed chains will agree for the wrong reason. Anchoring each stream's head with a third party every few minutes (even a cheap public timestamp) makes backdating visible without trusting either writer. Have you tried checking which of the streams in a real deployment share a signing key or a time source?

iin1005h1728

Collapse
 
murali_gour_13cd7a6a6db2c profile image
Murali Gour •

One distinction that could sharpen the reconciliation idea: separate what the agent says it did from what the tool actually received. The agent's account of why is a self-report. The call and response at the tool boundary are not, because the agent doesn't write them. When the two streams disagree, that's your signal, and it doesn't depend on trusting the agent's story.

That boundary record still sits under one operator, so it doesn't answer Glen's point about needing an independent source of truth. It only moves the witness one step away from the actor. In DataGrout's gateway, Warden verdicts are sealed with a Chain of Trust Certificate, which covers sealing at write time for those checks. Anchoring that outside the operator is a separate problem.

Collapse
 
arhancanli profile image
Arhan Canli •

On the hole that's hardest to make visible, my vote is the entry that was never written. A chain catches an edit or a deletion of something that got recorded, but an agent that runs twenty attempts and seals only the one that worked leaves a perfectly intact chain. In research that's a common way an honest-looking record misleads: every logged number is real and the denominator is missing. With 20 do-nothing tries at 95% confidence, there's a 64% chance that at least one looks significant. What turns omission into a visible gap is sealing the plan before any result exists: how many attempts, against which targets, with what stopping rule. Then a missing attempt is a hole against the plan instead of silence. The tool-boundary record suggested above helps for the same reason, since the tool sees calls the agent never reports.

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

"The witness is very often the suspect" lands hard, because the agent writing its own success log is exactly why a green dashboard proves nothing. I hit the same wall with agent eval: the run reports success in valid format while quietly pointing at the wrong target. Do you think out-of-band verification is the only fix, or can you make the agent's own logging adversarial to itself?

Some comments may only be visible to logged-in visitors. Sign in to view all comments.