DEV Community

Cover image for I built the flight data recorder for AI agents.
Sam LABBE
Sam LABBE

Posted on Edited on

I built the flight data recorder for AI agents.

Or: why I spent my evenings making AI-agent journals impossible to rewrite.


One Tuesday, at 2:37 PM

A disgruntled employee opens the database behind an AI product and
edits one row. The minutes an agent produced yesterday — the ones a client
is about to receive — now say something else. No alert. No trace. The
dashboard shows a pristine history, because the history has been
rewritten
.

This isn't a thriller plot. It is the default state of every AI-agent
product shipping today: their logs live in mutable databases, and nothing
exists to prove, after the fact, what the agent actually decided
.

I looked for the tool that closes this gap. It didn't exist. So I built it.
It's called NoireBox, it's open source (MIT), and this post is about
what it does, how it does it, and why I believe it's a missing piece.

→ https://github.com/slabbdev/noirebox

The founding question

After fifteen years of backend engineering — insurance, renewable energy,
real estate — in systems where every decision must be defensible years
later, I asked the question that started everything:

If someone contests a decision made by an AI agent tomorrow, who can
prove it?

Not "who can tell the story". Not "where is it logged". Who can prove it
— to a client, a lawyer, an auditor, a regulator — without trusting the
vendor, the host, or the administrator?

The answer, everywhere, was: nobody. Observability tooling (Langfuse,
LangSmith, Helicone…) is great for debugging what happened — but its
logs stay mutable, and nothing is exportable as evidence. Debugging and
proving are different jobs.

The principle: a proof, not a promise

NoireBox is a tamper-evident journal for AI agents — a flight data
recorder, like the ones in aircraft. Every event (model call, output,
decision, incident) goes through four mechanisms:

1. Hash chaining. Each event's SHA-256 fingerprint includes the previous
event's. Modify, insert, or delete anything and the chain breaks — and the
verifier tells you which event was tampered with. A numbered notebook:
you can't tear out a page without anyone noticing.

2. Signatures. Every fingerprint is signed with Ed25519, using a key
that never leaves the server (chmod 600). Regenerating the whole chain
without the key? The forgery is visible on sight.

3. The outside witness. The detail that kills tampering at the source:
NoireBox has an external timestamp authority (RFC 3161 — the same
protocol notaries and audit firms use) sign the chain head's fingerprint at
time T. Rewriting history afterwards becomes arithmetically impossible:
the old token no longer covers the new head. A journal that timestamps
itself proves nothing — that's the suspect writing its own police report.

4. The third-party-verifiable export. One file, plus a standalone
verifier that anyone runs on their own machine, offline, with no
credentials:

$ python verifier/verifier.py export.json
[✓] INTACT — 1260 events verified, attestation valid, 3 anchor tokens (1 against pinned roots), head of chain: 521de8d31240d0d0…
Enter fullscreen mode Exit fullscreen mode

Output from the journal running on my machine as I write these lines —
not an illustration.

And on that Tuesday at 2:37 PM, the same verifier returns a different
verdict:

[3] An attacker rewrites event 2: the minutes now read "REFUSE".
[✗] DETECTED — event 2: invalid hash (content was modified)
Enter fullscreen mode Exit fullscreen mode

(Real output from the repo's tamper demo — not an illustration.)

NoireBox never asks you to trust it. It hands you the dossier, and anyone
recomputes the truth themselves.

What about a whole fleet? One seal, a thousand boxes

Each box seals its own chain head individually. But what if you run 1,200
agents? Aggregated anchoring: the chain heads of N boxes form a Merkle
tree
whose root receives a single TSA seal. One seal for the whole
fleet — and each box proves it took part with ~log₂(N) hashes (11 for
1,200), verified offline.

That's the Certificate Transparency model — the system that made HTTPS
certificates auditable worldwide — applied to AI-agent decisions. To my
knowledge, nobody does this.

$ make demo-fleet
[2] Merkle tree: 3 leaves, root 7f2cf1a1e7b9879d…
    cr-reunion         proof: 2 hashes → ✓ covered
    support-juridique  proof: 2 hashes → ✓ covered
    scoring-credit     proof: 2 hashes → ✓ covered
[5] support-juridique regenerates its journal: content rewritten, re-chained cleanly.
[✗] DETECTED — the new head is not covered by the fleet seal.
    old head covered: yes; new one: NO
Enter fullscreen mode Exit fullscreen mode

(The root is different on every run — it commits to the exact state of
every chain at sealing time. That's the point.)

The most important part: the hub doing the aggregation cannot cheat. It can
omit a box (a detectable silence), but it can neither rewrite a sealed
batch nor include a fake head — and it journals the full tree into its
own
NoireBox journal. Every member verifies locally, without
contacting anyone. The arithmetic decides.

The guardrail is a plugin (keep yours)

Already running Lakera, Llama Guard, your own LLM-judge, or your own
regexes? Keep them. Prevention is a fungible layer — everyone has their
own, and it will keep changing. Proof is the universal layer.

Journal your solution's verdicts with a single POST /api/v1/events
(type: "incident") and its catches become tamper-evident and
third-party-verifiable, instead of ending up in rewritable app logs. The
bundled detector (regex + micro-model, French and English, ~250 KB per
language) is a working example of the plugin contract, not an
obligation.

The numbers — all reproducible with one command

Metric Value Reproduce with
Sealing one event (hash + signature + commit) 0.18 ms make bench
Full verification of 100 events ~46 ms make bench
Bundled detectors, FR + EN 293 + 243 KB make train + make train-en
Held-out attack sentences (never seen in training) 14/14 tests/test_ml_guardrail.py
Test suite 126 green make test

Performance is boring by design, as it should be: sealing an event
(hash + signature + commit) costs a fraction of a millisecond, and
verifying a 100-event export takes about a tenth of a second on a laptop.
The bottleneck is SHA-256 and Ed25519 — not clever code.

I imposed one rule on myself: no number that can't be reproduced with one
command.
A metric we can't reproduce is unknown — not a rounder number.
That's also why this post contains no "78M+ events", no customer logos, no
testimonials.

What it is not

  • Not a blockchain. One issuer, one verifier. Blockchains solve a problem we don't have — consensus among strangers — at a complexity price we refuse to pay.
  • Not a certification. A building block that feeds your audits; the exact perimeter is written down in the threat model.
  • Not an LLM. A 250 KB specialist detector that sorts sentences into 5 classes, deterministic and testable — not a black box predicting the next token.

Why now

The EU AI Act (article 12)
requires automatic event logging for high-risk systems. GDPR (art. 5(2),
30) puts the burden of proof on whoever processes the data. ISO 42001
asks for traceability of AI decisions.

European vendors will have to prove — not promise — what their agents
did. Today, almost none of them can. NoireBox is a sovereign building block:
self-hosted, zero telemetry, keys stay yours. Built in the Vosges mountains
🇫🇷 — yes, proof infrastructure can be built from the mountains.

Try it in three commands — or one with Docker 🐳

$ git clone https://github.com/slabbdev/noirebox && cd noirebox
$ ./start.sh          # 126 green tests + API on :8768 — /docs is live
$ make demo           # the model → the journal → the auditor → the attacker
Enter fullscreen mode Exit fullscreen mode

Or from PyPI — the ML engine ships inside the wheel:

$ pip install noirebox
$ noirebox serve      # API + OpenAPI docs on :8768/docs
Enter fullscreen mode Exit fullscreen mode

Or with Docker — no Python needed:

docker run -p 8768:8768 ghcr.io/slabbdev/noirebox:latest
Enter fullscreen mode Exit fullscreen mode

API + OpenAPI docs on :8768/docs — data persists in ./data. The image is
public on GHCR: pull it anonymously, run it, verify it.

FastAPI, 10 documented routes, a Python SDK, an MCP server (4 tools to plug
your agents in), a DPO-ready PDF export, and a self-hosted TSA (make tsa,
OpenSSL, $0, offline) — because a witness you don't control is a witness
less.

What's next

Recently shipped: a read-only supervision dashboard (GET /dashboard) with
an activity heatmap, and a reconciliation plugin born from a community
discussion — decision/outcome pairing by correlation key, findings sealed
into the journal (issue #3).
Next on the roadmap: a per-writer identity field, HSM/KMS-backed keys. The
core is MIT, forever — verification included.

If this resonates, here's what would help the most:

  • ⭐ A star on the repo — it sounds silly, but it's what makes a project exist
  • 🕵️ Try to break the proof. This is a project whose product is resistance to tampering: the best compliment is an issue describing an attack
  • 🔀 Fork it, PR it, question it — the code is deliberately small (one journal, one plugin, one verifier) so it can be read end to end

And if you know a DPO, a CISO, or an AI vendor preparing for article 12,
send them this post — that's exactly who this black box is for.


NoireBox — every AI decision, sealed forever.

Built solo, in the Vosges mountains, between two missions. Code, threat
model and specs:

github.com/slabbdev/noirebox

Top comments (7)

Collapse
 
axiru profile image
Axiru •

Sealing the decision and not just the outcome is the part most logging skips. For a refund or payout tool there are two events to seal: the policy decision (allow, hold or deny, with the reason code and the policy version) and the provider response after the money moved.

How do you handle the case where they disagree? An allow that never reached Stripe because the process died, or a charge that went through with no sealed decision in front of it. Does NoireBox flag the missing half, or is reconciliation left to the caller?

Collapse
 
slabb profile image
Sam LABBE • • Edited

This is exactly the right question — decision vs outcome is where most logging quietly lies.

What NoireBox guarantees today: the integrator seals both events (your policy_decision — allow/hold/deny, reason code, policy version — then the provider_response), and the journal proves the order and the integrity: the decision's hash sits inside the outcome's prev_hash chain, with millisecond timestamps and signatures. After the fact, nobody can claim "the charge ran first" or "that decision was added later" — the sequence is signed and anchored.

What NoireBox deliberately does not do: know what a "missing half" is. The journal is domain-blind by design — an event is a type and a payload, which is why the same box works for payouts, medical AI or legal. So, your two cases:

  • Process dies after sealing the decision, before Stripe: the journal pins it exactly — decision sealed at T, no outcome recorded. It doesn't fabricate knowledge of what happened next, but the gap itself is provable evidence: "here is what the system knew at the moment it stopped." That's the difference between reconstructing an incident from a mutable log vs from a flight recorder.
  • A charge with no sealed decision in front of it: if the caller skips journaling, the journal honestly shows an outcome with nothing before it — detectable by sequence order + a correlation id in the payloads (a one-line query over the export) — but NoireBox won't flag it on its own today. Reconciliation is caller-side.

That last point is a fair critique, and it's on the roadmap: a reconciliation plugin — a scheduled check over the journal that encodes your invariants ("every decision must have an outcome within N minutes") and journals its findings as reconciliation events. Same plugin pattern as the guardrail: the journal stays domain-blind, the domain knowledge stays pluggable.

If you're building a payout flow, the two-event pattern with a correlation id works today — and I'd genuinely like the invariant case sketched as an issue. It would shape the reconciliation plugin, and you'd be credited for it.

Collapse
 
slabb profile image
Sam LABBE •

Opened an issue to track this → github.com/slabbdev/noirebox/issues/3. Thanks for the sharpest question of the thread.

Thread Thread
 
naveen_alavilli profile image
Naveen Alavilli •

This solves a problem worth solving well — tamper-evidence after the write is a real gap and hash-chaining plus an external TSA closes it properly. The boundary worth naming for anyone evaluating it: NoireBox proves nobody edited event N after it was sealed. It can't prove event N was true when it was sealed. If the integrator's own instrumentation misrepresents what the agent actually did — logs "approved by policy X" when the check silently no-opped, say — the chain still verifies clean, because the lie shipped at write time, not after. That's not a flaw in the design, tamper-evidence and correctness-at-source are genuinely different problems, but it's worth being explicit that this tool answers "was the record altered," not "was the record accurate." Both matter for the EU AI Act use case, and only one of them is what a hash chain can give you.

Thread Thread
 
slabb profile image
Sam LABBE •

@naveen_alavilli You've drawn the boundary exactly where we'd draw it — and it deserves your precision: NoireBox answers "was the record altered?" and "when did this claim exist?", never "was the claim true when written?". A lie sealed at write time verifies clean forever; the chain attests a lie with exactly the same force it would attest the truth.

Two things the chain still changes in your "approved by policy X" scenario, though:

1. The lie is pinned. Once sealed, the story can't be quietly rewritten when the audit starts — "that's not what our logs said" is off the table, and the third-party timestamps (RFC 3161 / Bitcoin — ADR 006 + 009) fix when the claim existed. The chain doesn't detect the lie; it makes the lie non-retractable and attributable to a key.

2. Write-time lies become falsifiable when the sources are separated. That's what the reconciliation layer (issue #3, v0.1, shipped) is for: have the policy engine seal its own decision events, the instrumentation seal its claims, and noirebox reconcile correlate them — an outcome ("approved, money moved") with no sealed policy decision in front of it surfaces as orphan_outcome, and noirebox reconcile --fail-on-findings gates CI on it (demo_payout.py runs exactly those two failure cases). A hash chain can't validate truth, but it makes claims from independent writers cross-checkable against each other after the fact — fooling it now requires lying at both sources.

One caveat in the spirit of your comment: one instance = one key pair, so "independent writers" today means independent processes posting their own events, not independent keys — sealer identity is the natural next step. And you're right that the threat model should carry this boundary explicitly rather than leave it implied — PR #19 adds exactly that line. The most useful kind of review — thanks.

Collapse
 
slabb profile image
Sam LABBE • • Edited

UPDATE 🐳 — it now runs in one command.

No clone, no Python, no setup:

docker run -p 8768:8768 ghcr.io/noirebox/noirebox:latest
Enter fullscreen mode Exit fullscreen mode

API + OpenAPI docs on :8768/docs, the ML guardrail ships inside the image, and every incident it catches is sealed in the journal before you even open a browser. Pull it anonymously, run it, verify it — that's the whole point of a flight recorder.

Small print, as always: the image is public, the tests are green, the numbers are reproducible. ☕ If the project helps: buymeacoffee.com/samlabbe

Some comments may only be visible to logged-in visitors. Sign in to view all comments.