DEV Community

Cover image for A Field Guide to AI Documentation: Model Cards, Eval Reports, Agent Cards, and More

A Field Guide to AI Documentation: Model Cards, Eval Reports, Agent Cards, and More

James Anderson on September 26, 2026

Every developer learns the same documentation types. The README. The API reference. Code comments. Maybe an architecture doc if the team is discipl...
Collapse
 
aniketsahu141 profile image
Aniket Sahu •

This is an interesting shift in how we think about documentation. AI systems introduce context, decisions, data, and behavior that aren’t always visible in the code itself. Documentation is no longer just about helping the next developer understand the system, it’s increasingly about making AI systems explainable, traceable, and accountable.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — documentation stopped being "help the next developer read the code" and became "prove the system is explainable, traceable, and accountable" when the behavior stopped living in the code at all.

Collapse
 
dexoryn profile image
Dexoryn •

The documentation drift point is the one I’d worry about most.
An eval report can be accurate today and misleading after a model or prompt change.
Should these docs be generated and versioned alongside the system, rather than maintained manually?

Collapse
 
james_anderson_h profile image
James Anderson •

Yes — and it's the trap the thread converged on: a hand-maintained eval report doesn't just age, it ages confidently, vouching for a state that's gone while looking maintained. Generate and version them as build artifacts — same pipeline, config-in/markdown-out, committed with the code, CI diffs on every change so a hand-edit becomes a build failure, not silent drift.

Collapse
 
contentclips_st profile image
ContentClips •

Nice taxonomy — the four questions framing (trained on what / measured how / allowed what / actually did what) is a good litmus test: if a doc can't be mapped to one of them, it's probably not a trust artifact. One hard-won lesson from maintaining these: model cards and eval reports rot faster than READMEs. A hand-written card drifts from the model it describes the moment a retrain lands. What helped was treating them as build artifacts rather than documents — generated from the same pipeline that produces the model (config + eval results in, markdown out), committed alongside the weights, with a stale flag in CI when inputs change. Same for agent cards: the "what did it actually do" section is only trustworthy if it is derived from run logs, not written from memory. Hand-editing trust artifacts quietly defeats their purpose.

Collapse
 
james_anderson_h profile image
James Anderson •

The rot point is the one I under-weighted. A drifted trust artifact is worse than none — it confidently vouches for a state that no longer exists, and everyone trusts it because it looks maintained. That's silent-success in the docs layer. Generating them as build artifacts from the pipeline is the only real fix; the moment a human hand-edits, the drift reopens.

Collapse
 
contentclips_st profile image
ContentClips •

Agreed — and the one thing generation alone doesn't fix is freshness at consumption time. A build artifact can still confidently vouch for a stale state if the pipeline runs on a schedule instead of on change. Two things that help in practice: stamp every generated artifact with provenance (model version, eval-set hash, generation timestamp) so staleness stays visible downstream even if the artifact escapes, and add a CI freshness gate that regenerates and diffs — if the committed artifact doesn't match a fresh run, the build fails, exactly like a lint check. That diff doubles as a changelog too: reviewers see what moved between versions instead of re-reading a report that looks identical. And it flips hand-edits from quiet drift into a build failure, which is the only way "generate, never hand-edit" survives contact with a deadline.

Thread Thread
 
james_anderson_h profile image
James Anderson •

Provenance stamping plus a CI freshness gate is the piece the whole thread was missing — generation closes the drift, but only a regenerate-and-diff-on-every-change actually enforces it. And flipping hand-edits from silent drift into a hard build failure is the only thing that makes "never hand-edit" survive a deadline, because discipline alone never does. The diff-as-changelog bonus is the part I'll steal — reviewers seeing what moved beats re-reading a report that looks identical.

Collapse
 
micheypico profile image
Micheal Heypico •

This matches what we see operating a model-routing layer (32 models, one key at heypico.ai): the deterministic scaffolding around the LLM is what makes multi-model setups viable. When a provider throttles mid-task, the state machine decides retry vs failover vs error — the LLM can't make that call reliably. Debugging a 'flaky agent' is usually debugging a missing state machine around a fine model.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — "debugging a flaky agent is usually debugging a missing state machine around a fine model" is the sharpest version of the whole idea. Retry-vs-failover-vs-error on a mid-task throttle is a deterministic decision, and handing it to the LLM is where the flakiness comes from — the model was never the problem, the missing scaffolding was.

Collapse
 
contentclips_st profile image
ContentClips •

Agreed - and one thing that makes the drift detectable instead of just likely: every generated artifact embeds the source commit hash and a generation timestamp, so the looks-maintained-but-isnt case becomes greppable. The artifact answers the staleness question itself instead of trusting its own confidence, and a hand-edit then stands out as the anomaly, not the default. One extra line in the generator, and is-this-stale becomes a machine answer.

Collapse
 
james_anderson_h profile image
James Anderson •

Embedding the source commit hash is the move — it makes the artifact answer "am I stale?" itself instead of vouching for its own confidence, so a hand-edit becomes the visible anomaly instead of the silent default. One line in the generator turns a trust question into a grep.

Collapse
 
theagentloop profile image
The Agent Loop •

The four-question split is the right one, and \"what did it actually do\" is the one that's hardest to answer in practice. Model cards and datasheets describe intent and inputs; the artifact I keep reaching for after an incident is the run record: which tool was called, with what arguments, and what it cost. It's the only document on this map that can't be written in advance, because it's a record rather than a description.

That's also where I think the map is thinnest. Everything here describes the system at rest or at design time, and nothing here answers \"what happened last Tuesday\" at per-tool granularity. When an agent bill or an incident gets questioned, a total with no attribution tells you a spike happened and nothing about what to change, so we've started treating that per-run ledger as a doc type of its own.

Which of these does your team actually produce by hand versus generate? My guess is model cards are hand-written, eval reports are half-generated, and agent cards land in between, but I'd rather hear it from someone who has had to file them.

Collapse
 
james_anderson_h profile image
James Anderson •

You've named the sharpest gap on the map: everything else describes the system, but the run record is the only thing that is a record — written after the fact, per-run, and the one artifact you actually reach for after an incident. "A total with no attribution tells you a spike happened and nothing about what to change" is exactly why the per-tool ledger deserves to be its own doc type, not a footnote under audit trails — I under-drew it, and you're right to pull it out. On the hand-vs-generated question: my honest read matches yours — model cards are hand-written because they encode intent only a human holds, eval reports are half-generated (the numbers auto, the methodology and caveats by hand), and agent cards sit in between. But I'd genuinely rather hear from people filing them at scale, because I suspect the answer drifts toward "generated" faster than anyone admits — which quietly reintroduces the exact trust problem the docs existed to solve.