DEV Community

Open Human
Open Human

Posted on

Constitutional Engineering: What Two Days of a Three-Copy Word List Taught Me About Agent Governance

Put Governance Rules Where They Can Bite

The word list lived in three places: the generator, the publisher, and the manual review queue. They were supposed to be identical. Adding a new AI-flavor phrase meant editing all three files, and every time we missed one, something slipped through a gate. First miss was the publisher. Second miss was the generator. Third miss was the human review queue — a person approved an article that should have been flagged, because the copy of the list in front of them was stale. A diff across the three files after that third incident took a few seconds and made the case better than any slides: the phrase we'd added was in two files, and the one it was missing from was the gate that had just approved the article. Three consecutive misses at three different gates, same root cause: one rule, three copies.

We spent two days consolidating the lists into a single source of truth. While we were in there, we split the words into two tiers. Blocking words stop the pipeline. Prompting words add a note for the human reviewer. The old one-tier list was a trap: add a high-frequency connector to fight AI-sounding prose and every normal article that used it got flagged. "However" is not a crime. "But" is not a crime. But in a one-tier list, a weak signal and a strong signal carried the same weight. The tiering fixed that. Now adding a word starts with a question: which tier? And the answer, more often than not, is prompting.

That consolidation changed how I think about rules in the multi-agent memory systems we've been building. In the MCP universe, every agent has a memory socket, a tool belt, and a mandate to get things done. The failure mode is memory poisoning via shared context: an agent misreads stale data, a rogue tool call overwrites a critical state, or a delegated sub-agent inherits a memory scope it should never see. A malicious attacker is a rare event. A planning agent that read a stale entry, made a plausible downstream call, and watched that call overwrite a critical state — that's a Tuesday. The damage shows up as hours of tracing why a recommendation chain collapsed.

The word list taught us that a rule kept in one file as documentation is a suggestion. A rule kept in three files is three suggestions. We keep calling it constitutional engineering: governance rules embedded in the protocol itself, so the system physically cannot take a prohibited action. A YAML constitution is a readme. Enforcement has to live at the point of mutation.

The first thing we put in place was fixed-point verification on memory writes. Every MCP tool call that mutates shared memory carries a deterministic hash of the entire conversational context it was derived from. The memory layer rejects any write whose hash mismatch suggests the agent hallucinated or omitted a conflicting fact. That is how you prove causality in shared state.

Our first attempt placed the verification in the orchestration wrapper — the Python layer around the tool call. It held for about a week. Then an agent called the memory socket directly, the wrapper never ran, and the hash was computed over a context snapshot we didn't recognize. The guard has to live inside the thing it guards. The verification now sits in the memory kernel; the tool definition requires the hash field, the kernel computes it independently, and it rejects before persisting.

The cost is real. Shared-memory writes got slower, around 40% latency increase on full-context hashing. We tried full verification on every shared namespace, and the latency pushed teams to route around it — writing "context summaries" to a side cache that was never verified. That is how you get two sources of truth that disagree. We tried relaxing verification for a "low-stakes" namespace, and a rogue tool call wrote junk there that a planning agent used as fact for three days. Current position: verify all cross-agent writes, skip verification only for the agent's own scratch memory. It's a compromise we haven't fully validated, but it's holding.

The second pattern came out of that same consolidation: the frozen memory pattern. Memory entries older than X days cannot be referenced by a tool call unless a human explicitly re-activates them. A provenance fence. It stops the plausible, dangerous habit of a planning agent relying on volatile knowledge that a sibling agent silently updated.

The first version used a single global X. Wrong. The planner's long-term context froze in a way that looked like a planner bug, and we burned a day before finding the namespace overlap. The fix was per-namespace freeze windows: decision logs freeze at 14 days, sensor readings at 5, an agent's own working output at 30. Every MCP tool response carries a version number on its memory reference. When a chain comes apart, the version number shows you which snapshot each hop used. It doesn't end the incident, but it cuts the search space to something a human can handle in one sitting.

One organizational detail bit us. The freeze window originally lived in runtime config. Someone on call pushed it to 60 days during an incident to quiet the alerts, and stale references came back. The freeze window now lives in the kernel binary. To change it, you ship a new build. Runtime tuning is how governance parameters get un-tuned.

The third piece is procedural segregation in MCP scopes. We don't give agents a full memory schema. Micro-scopes: read_recent_own_output, read_shared_decision_log, read_sensor_cache. Least privilege applied to the past — an agent cannot see what it cannot govern. Our first attempt scoped by role. The "planner" role could read everything because "it needs context," and it planned beautifully on memory fragments it had no authority to synthesize. Scopes now attach to the memory namespace. Role identity no longer grants memory access. Delegation re-derives scopes from the sub-task definition — a sub-agent inherits only the namespaces its parent was granted for that specific task, not the parent's full context. If the sub-task doesn't name a namespace, the sub-agent starts empty. This one curbed the most failures in practice. The "why did this agent recommend that" tickets dropped sharply after the change.

The trade-offs are honest and ugly. Byte-level verification makes the system slower. Procedural segregation means more MCP tools to build and maintain — we landed on twelve composable micro-scopes, and we tried consolidating them into a single read-all scope to cut the tooling burden. The visibility regression came right back. We keep returning to the same position: every limitation you accept on the agent side is a capability you remove from an attacker — or from an idiot agent, which is far more common in our logs.

The ecosystem gap is audit trails. There's no standard for replaying why a given memory item was written, verified, and authorized to survive. We need something like git bisect for memory lineage. Until then, we grep timestamps, compare snapshot hashes, and every incident review starts with "who wrote this entry and under what scope." Deprecating a bad memory rule without turning all agents into half-blind experts is a problem we haven't solved.

Back to the word list. If the old setup had the same architecture we now use for agent memory, the third miss would never have happened: one list, one entry point, tiers enforced where words are used. Recovery would have been minutes, not two days. Write your governance rules where they can bite. If a rule survives only as a static policy file

maref #ai #opensource #machinelearning

Top comments (0)