DEV Community

Cover image for A Field Guide to AI Documentation: Model Cards, Eval Reports, Agent Cards, and More
James Anderson
James Anderson

Posted on AI-assisted

A Field Guide to AI Documentation: Model Cards, Eval Reports, Agent Cards, and More

Docs rot faster than code

Every developer learns the same documentation types. The README. The API reference. Code comments. Maybe an architecture doc if the team is disciplined. That was the whole vocabulary for decades.

Then AI arrived and quietly created an entire new category of documentation — one nobody taught us to write, that isn't in any bootcamp or CS degree, and that you're increasingly expected to produce anyway. Sometimes because a teammate needs it. Sometimes because an auditor does. And, as of 2026, sometimes because it's the law.

This is a field guide to that new landscape. I'll walk through the real doc types, grouped by what they document, with what each one is, when you need it, and why it exists. But first, the one idea that ties all of them together — because once you see it, every one of these makes sense.

Why AI needed new documentation at all

Here's the thing that broke.

For our entire careers, documentation described deterministic behavior. A function takes an input and returns the same output every time. You could document exactly what it does, because it does exactly one thing. The code itself was, in a sense, the ultimate documentation of what would happen.

AI shattered that assumption. A generative model can produce a different output for the same input twice in a row. An agent can take a path nobody wrote. You cannot document exact behavior for a system that doesn't behave the same way twice — so documentation had to shift from "what it does" to a different set of questions:

  • What was it trained on?
  • How did we measure it?
  • What is it allowed to do?
  • What did it actually do?

Every doc type below is an answer to one of those four questions. They aren't paperwork — they're trust artifacts. In a deterministic world, trust came for free from the code. In a probabilistic world, trust has to be manufactured through documentation, because it's the only way anyone — a teammate, an auditor, a regulator, a user — can believe your AI does what you claim. Keep that in mind and the whole landscape organizes itself.


Group 1: Documenting the model and the data

These are the foundational ones — the docs that describe the ingredients of an AI system.

Model cards

What it is: A short, structured document describing a single trained model — its intended use, the data it was trained on, its evaluation results, its limitations, and its out-of-scope uses. Think of it as a nutrition label for a model: at a glance, what's in it and what it's safe for.

When you need it: Any time you ship, publish, or hand off a model someone else will rely on. Introduced by Mitchell et al. in 2019, model cards are now standard on model hubs like Hugging Face — and under the EU AI Act, they (or an equivalent) are a regulatory requirement for high-risk systems. If you fine-tune a model and give it to another team, they need a model card to use it responsibly.

Why it exists: Because "here's a model, good luck" is how you get someone deploying a system into a context it was never meant for. The model card answers what is this model, really, and where should it not be used — the questions the weights themselves can't tell you.

Datasheets for datasets

What it is: The data-side twin of the model card. A structured document describing a dataset — where it came from, how it was collected, its composition, its known biases, consent and licensing, and what it should and shouldn't be used for. Introduced by Gebru et al. (2018/2021).

When you need it: Whenever a dataset outlives the person who made it — which is always. If your model's behavior is ever questioned, the first question is "what was it trained on?", and a datasheet is the answer you'll wish you'd written at the time.

Why it exists: Because most AI failures trace back to the data, and data provenance is almost never obvious after the fact. A datasheet forces the dataset's creators to be intentional and honest about its origins and limits — before those limits become someone's incident.

System cards

What it is: A step up from a model card. Where a model card documents one trained model, a system card documents the complete AI system — often a pipeline of multiple models, plus architecture, safety evaluations, operational constraints, and deployment context. Crucially, a system card documents how the system behaves, not how it's used (that's the difference from traditional product docs).

When you need it: When your AI isn't one model but a stack — a router, a fine-tune, retrieval, chained components (which, in 2026, is most real systems). Frontier labs publish system cards for their big models (the GPT system cards are the canonical examples) precisely because "the model" is really a system.

Why it exists: Because modern AI is rarely a single model, and documenting only the model misses everything that emerges from the combination. The system card is where a forward-deployment engineer figures out whether the whole thing fits their use case.


Group 2: Documenting the evaluation

This is the newest and most underrated genre — and arguably the most important, because it's where trust is actually earned or faked.

Eval reports (eval factsheets)

What it is: Documentation of how you evaluated an AI system and what you found — the test set, the methodology, the assumptions behind the metrics, and how confident you should be in the results. "Eval factsheets" formalize this with the same rigor datasheets brought to data.

When you need it: Any time you make a claim about how well an AI performs — which is any time you ship one. Especially critical for agents and anything customer-facing.

Why it exists: Here's the sharp insight that justifies the whole genre. A model card gives you the score, but not how the score was obtained — and the "how" is the part that tells you whether to trust it. "95% accuracy" means nothing until you know on what data, under what conditions, with what definition of correct. Eval reports document the methodology, not just the number. A benchmark result with no eval report is a marketing claim, not evidence.

Regression evals and benchmark cards

What it is: Two close relatives. A regression eval is documentation of the tests you re-run to make sure a model or prompt change didn't quietly break behavior that used to work (the AI-era version of a regression test suite). A benchmark card documents the benchmark itself — what it actually measures, its limitations, how to read its scores — so people stop treating a leaderboard number as gospel.

When you need it: Regression evals the moment you're changing a model or prompt in production. Benchmark cards whenever you publish or cite a benchmark.

Why it exists: Because AI systems drift, and "it worked before" isn't verifiable without a documented, repeatable eval — and because raw benchmark numbers, undocumented, mislead more than they inform.


Group 3: Documenting the agents

Brand new territory, and the most relevant to anyone building right now. Until recently there was no analog to model cards for agents. In 2026, that changed.

Agent cards

What it is: A 2026 documentation standard for operational AI agents. Where a model card describes a model, an agent card captures an agent's operational attributes: its roles, its memory taxonomy, its tool integrations, its communication protocols, its monitoring hooks, its governance scope, and its evaluation metrics.

When you need it: When you deploy an agent that acts — calls tools, touches data, runs unattended. Anyone who has to operate, audit, or trust that agent needs to know what it can reach and how it's monitored.

Why it exists: Because an agent is a fundamentally different thing to document than a model — it acts, it has memory, it uses tools, it runs over time. Model cards and datasheets couldn't capture any of that, so a new artifact had to fill the gap. The agent card is how you make an acting system transparent, comparable, and auditable.

Policy cards

What it is: Machine-readable runtime governance for autonomous agents — documentation of what an agent is allowed to do, in a form that can actually be enforced while it runs, not just read by a human afterward.

When you need it: For any agent with real permissions — one that can send, spend, delete, or change things. This is the formalized version of a principle worth repeating: for a system that acts unpredictably, you can't document what it will do, so you document what it's allowed to do.

Why it exists: Because with agents, the interesting failure isn't "wrong output," it's "unsanctioned action." A policy card turns the boundaries into an artifact — ideally one the runtime checks — so an agent's permissions are explicit, auditable, and enforced instead of living in someone's head.

AGENTS.md / CLAUDE.md

What it is: The practical, in-the-repo instruction files that coding agents actually read — project conventions, constraints, architecture notes, and rules written for the AI to follow while working in your codebase.

When you need it: The moment you let an AI agent work in your repo. Without one, the agent improvises conventions; with one, it follows yours.

Why it exists: Because the agent is now a reader of your documentation — and it will faithfully do the wrong thing if you never wrote down the right thing. This is the doc type where "documentation" and "programming the AI" quietly merge.


Group 4: Documenting for governance and compliance

Rising fast, driven by regulation — and increasingly not optional.

Audit trails

What it is: An immutable, append-only record of every consequential model or agent decision in a production system — used as evidence for compliance, audits, and incident investigation.

When you need it: Any production AI that makes decisions with consequences. When something goes wrong (and it will), the audit trail is how you reconstruct what actually happened.

Why it exists: Because for a system that acts, the log is the documentation of what it did — and "what did it actually do?" is unanswerable without a durable, tamper-evident record. It's the fourth of our four questions, made concrete.

Compliance cards, AI cards, and transparency reports

What it is: A family of governance documents — compliance cards / AI cards (machine-readable risk and compliance documentation inspired by the EU AI Act), use case cards (documenting a specific deployment's intended use and risk), and transparency reports (public-facing accounts of what a system does and its limitations).

When you need it: When you operate under a regulatory regime — which is a rapidly growing "when." The EU AI Act mandates documentation for high-risk systems, and Singapore's Model AI Governance Framework for Agentic AI (launched January 2026) is the first framework explicitly requiring autonomous-AI documentation: risk assessments, limits on agent authority, human accountability at critical decision points.

Why it exists: Because AI documentation stopped being a nice-to-have and became a legal artifact. Governance without documentation is just a promise, and regulators have stopped accepting promises.


The idea that ties it all together

Step back and look at the whole map, and the pattern is unmistakable.

Model cards and datasheets answer what was it trained on? Eval reports answer how did we measure it? Agent cards and policy cards answer what is it allowed to do? Audit trails answer what did it actually do? Every single one of these documents exists because AI broke the thing that used to make documentation easy: determinism. When a system won't behave the same way twice, you can't document its behavior — so you document its ingredients, its measurement, its boundaries, and its history instead.

That's why these aren't bureaucracy, even though they can feel like it. They're trust artifacts. In the old world, the code told you what would happen, so trust was free. In this world, trust has to be built, deliberately, in writing — because a model card is how a teammate trusts your fine-tune, an eval report is how a reviewer trusts your accuracy claim, a policy card is how an operator trusts your agent near production, and an audit trail is how an investigator trusts your account of an incident.

The takeaway

Nobody taught us to write these. There was no reason to — until the systems we build stopped being predictable, and "read the code" stopped being a sufficient answer to "what does this do?"

The developers who thrive in this era won't just be the ones who can build AI systems. They'll be the ones who can make those systems trustworthy to everyone else — and increasingly, that's a documentation skill. It's an oddly human one, too: judgment about what to disclose, honesty about how you measured, clarity about where you drew the boundary. The machine can help you write these. It can't decide what belongs in them. That part is still yours.

So the next time you ship something with a model in it, ask the question we were never trained to ask: not just "does it work?" but "could anyone else tell that it works, and where it doesn't, from what I wrote down?" That question, and the docs that answer it, are the quiet infrastructure the whole AI era is going to run on.


Which of these have you actually had to write — and, honestly, which did you not know existed until just now? I'd genuinely like to know where the real practice is, versus where the standards say it should be. And if there's a doc type you're using that I left off the map, add it below.


Sources & further reading: Model Cards (Mitchell et al., 2019); Datasheets for Datasets (Gebru et al., 2018/2021); System Cards (frontier-lab system cards, e.g. OpenAI's GPT system cards); Eval Factsheets and BenchmarkCards (2025–2026 evaluation-documentation research); Agent Cards and Policy Cards (2025–2026 agent-documentation standards); and governance frameworks including the EU AI Act and Singapore's Model AI Governance Framework for Agentic AI (Jan 2026). Standards in this space are evolving quickly — treat specifics as current-as-of-writing and check primary sources for the latest.

Top comments (14)

Collapse
 
aniketsahu141 profile image
Aniket Sahu •

This is an interesting shift in how we think about documentation. AI systems introduce context, decisions, data, and behavior that aren’t always visible in the code itself. Documentation is no longer just about helping the next developer understand the system, it’s increasingly about making AI systems explainable, traceable, and accountable.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — documentation stopped being "help the next developer read the code" and became "prove the system is explainable, traceable, and accountable" when the behavior stopped living in the code at all.

Collapse
 
dexoryn profile image
Dexoryn •

The documentation drift point is the one I’d worry about most.
An eval report can be accurate today and misleading after a model or prompt change.
Should these docs be generated and versioned alongside the system, rather than maintained manually?

Collapse
 
james_anderson_h profile image
James Anderson •

Yes — and it's the trap the thread converged on: a hand-maintained eval report doesn't just age, it ages confidently, vouching for a state that's gone while looking maintained. Generate and version them as build artifacts — same pipeline, config-in/markdown-out, committed with the code, CI diffs on every change so a hand-edit becomes a build failure, not silent drift.

Collapse
 
contentclips_st profile image
ContentClips •

Nice taxonomy — the four questions framing (trained on what / measured how / allowed what / actually did what) is a good litmus test: if a doc can't be mapped to one of them, it's probably not a trust artifact. One hard-won lesson from maintaining these: model cards and eval reports rot faster than READMEs. A hand-written card drifts from the model it describes the moment a retrain lands. What helped was treating them as build artifacts rather than documents — generated from the same pipeline that produces the model (config + eval results in, markdown out), committed alongside the weights, with a stale flag in CI when inputs change. Same for agent cards: the "what did it actually do" section is only trustworthy if it is derived from run logs, not written from memory. Hand-editing trust artifacts quietly defeats their purpose.

Collapse
 
james_anderson_h profile image
James Anderson •

The rot point is the one I under-weighted. A drifted trust artifact is worse than none — it confidently vouches for a state that no longer exists, and everyone trusts it because it looks maintained. That's silent-success in the docs layer. Generating them as build artifacts from the pipeline is the only real fix; the moment a human hand-edits, the drift reopens.

Collapse
 
contentclips_st profile image
ContentClips •

Agreed — and the one thing generation alone doesn't fix is freshness at consumption time. A build artifact can still confidently vouch for a stale state if the pipeline runs on a schedule instead of on change. Two things that help in practice: stamp every generated artifact with provenance (model version, eval-set hash, generation timestamp) so staleness stays visible downstream even if the artifact escapes, and add a CI freshness gate that regenerates and diffs — if the committed artifact doesn't match a fresh run, the build fails, exactly like a lint check. That diff doubles as a changelog too: reviewers see what moved between versions instead of re-reading a report that looks identical. And it flips hand-edits from quiet drift into a build failure, which is the only way "generate, never hand-edit" survives contact with a deadline.

Thread Thread
 
james_anderson_h profile image
James Anderson •

Provenance stamping plus a CI freshness gate is the piece the whole thread was missing — generation closes the drift, but only a regenerate-and-diff-on-every-change actually enforces it. And flipping hand-edits from silent drift into a hard build failure is the only thing that makes "never hand-edit" survive a deadline, because discipline alone never does. The diff-as-changelog bonus is the part I'll steal — reviewers seeing what moved beats re-reading a report that looks identical.

Collapse
 
micheypico profile image
Micheal Heypico •

This matches what we see operating a model-routing layer (32 models, one key at heypico.ai): the deterministic scaffolding around the LLM is what makes multi-model setups viable. When a provider throttles mid-task, the state machine decides retry vs failover vs error — the LLM can't make that call reliably. Debugging a 'flaky agent' is usually debugging a missing state machine around a fine model.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — "debugging a flaky agent is usually debugging a missing state machine around a fine model" is the sharpest version of the whole idea. Retry-vs-failover-vs-error on a mid-task throttle is a deterministic decision, and handing it to the LLM is where the flakiness comes from — the model was never the problem, the missing scaffolding was.

Collapse
 
contentclips_st profile image
ContentClips •

Agreed - and one thing that makes the drift detectable instead of just likely: every generated artifact embeds the source commit hash and a generation timestamp, so the looks-maintained-but-isnt case becomes greppable. The artifact answers the staleness question itself instead of trusting its own confidence, and a hand-edit then stands out as the anomaly, not the default. One extra line in the generator, and is-this-stale becomes a machine answer.

Collapse
 
james_anderson_h profile image
James Anderson •

Embedding the source commit hash is the move — it makes the artifact answer "am I stale?" itself instead of vouching for its own confidence, so a hand-edit becomes the visible anomaly instead of the silent default. One line in the generator turns a trust question into a grep.

Collapse
 
theagentloop profile image
The Agent Loop •

The four-question split is the right one, and \"what did it actually do\" is the one that's hardest to answer in practice. Model cards and datasheets describe intent and inputs; the artifact I keep reaching for after an incident is the run record: which tool was called, with what arguments, and what it cost. It's the only document on this map that can't be written in advance, because it's a record rather than a description.

That's also where I think the map is thinnest. Everything here describes the system at rest or at design time, and nothing here answers \"what happened last Tuesday\" at per-tool granularity. When an agent bill or an incident gets questioned, a total with no attribution tells you a spike happened and nothing about what to change, so we've started treating that per-run ledger as a doc type of its own.

Which of these does your team actually produce by hand versus generate? My guess is model cards are hand-written, eval reports are half-generated, and agent cards land in between, but I'd rather hear it from someone who has had to file them.

Collapse
 
james_anderson_h profile image
James Anderson •

You've named the sharpest gap on the map: everything else describes the system, but the run record is the only thing that is a record — written after the fact, per-run, and the one artifact you actually reach for after an incident. "A total with no attribution tells you a spike happened and nothing about what to change" is exactly why the per-tool ledger deserves to be its own doc type, not a footnote under audit trails — I under-drew it, and you're right to pull it out. On the hand-vs-generated question: my honest read matches yours — model cards are hand-written because they encode intent only a human holds, eval reports are half-generated (the numbers auto, the methodology and caveats by hand), and agent cards sit in between. But I'd genuinely rather hear from people filing them at scale, because I suspect the answer drifts toward "generated" faster than anyone admits — which quietly reintroduces the exact trust problem the docs existed to solve.