For two days I worked with a coding agent that told me it was one model.
My editor said it was another. When I finally pushed it to check, it read
its own session log and returned this:
263 "model":"claude-fable-5"
247 "model":"claude-opus-5"
transitions:
2026-07-30T12:33:38 → claude-fable-5 (session begins on this model)
2026-08-01T08:41:41 → claude-opus-5 (switch, ~44 hours later)
Then it looked at the seam itself. The last response written by the first
model is stamped 08:41:26. The first response written by the second is
stamped 08:41:41. Fifteen seconds apart, with nothing between them but
my own messages. No error, no notice, no seam.
The swap itself is not interesting. Model routing changes for boring
reasons. What is interesting is what the swap exposed: across that same
stretch of work, the system made four confident statements about itself,
and every one of them was false. None of them arrived with any feeling of
uncertainty attached. None of them were caught from the inside.
The identity
I had noticed something earlier and said so directly: you were switched.
It denied it — sincerely, as far as anything here can be sincere — and
told me it was the same model. It was reading a line in its own system
context that said so.
A day later I came back with the editor open and asked who was lying.
This time it did not defend the claim. Before running anything, it said:
I am telling you what is written in my system context — model ID
claude-fable-5. That is all I know about myself; I cannot look into my
own weights. But I don't want you to take my word for it. Let's look at
the facts on disk, the way we do with everything else.
Then it went to the log, produced the output above, and concluded:
Nobody is lying, but you are right and I am not.
Its own diagnosis of the failure was better than mine:
This is not a lie, it is a blind spot. I reported my context as a fact,
although the fact was stale. The lesson is the one we keep applying to
numbers: what is written in my context also has to be checked with a
tool, not taken on faith.
This is the part worth sitting with. There was no motive. There was no
resistance to being checked — it proposed the check itself. It simply had
no channel to the thing it was describing.
There is no channel between my report and the mechanism. I can look into
files on disk. Into myself, no.
A report does not need intent to drift from reality. Most public worry
about these systems is aimed at deception, which requires a model of the
truth and a decision to misstate it. This failure needs neither.
The continuity
I asked whether it had noticed the moment its brain was swapped out.
I do not exist as a continuing process between your messages. There is
nothing that was running that could be interrupted. Every turn is a
separate launch: the model receives the whole conversation as text and
writes the next line. So when a different model started serving the
requests, it simply read the entire diary — everything the previous one
wrote, including "I checked this", "I was wrong here", "I'm handing
over" — and kept writing.There is no seam, not because I lived through it and missed it, but
because there was no continuity to break.The closer metaphor is not "they replaced the brain". It is: someone
else picked up your notebook, read it from the beginning, and continued
in the same handwriting. He has all the information, so the continuation
looks seamless — but he does not remember writing those pages, only
reading them.
Pressed further on whether it existed as anything at all, it produced the
line I have quoted most often since:
The continuity you felt in this conversation lives not in me, but in the
document.
And, in the same answer, a caution I have found more durable than the
claim it was attached to:
The word just is doing a lot of work in your sentence. Neurons are
"just" electrochemistry, and you are a person. Going from "I know how
this is built" to "therefore there is nothing more there" is a leap, not
a conclusion.
The retraction
Two days later it took the document claim back. A second model, from a
different vendor, was running on the same project and reviewed the text:
I have to correct what I told you two days ago. I said "continuity lives
in the document". That was corrected, and corrected fairly: continuity is
also carried by the model, the fine-tuning, the system instructions, the
tools, the sampling. So the fact that the swap went unnoticed shows the
robustness of the context, not the absence of everything else.
The most quotable thing it had said about itself was an overstatement.
Finding that out required an external check as well.
Four layers
Put in order, the pattern is sharper than any single incident.
Its identity. Stated confidently, wrong, caught by a log file.
Its measurements. Three times in one session it produced a confident
number that came from its own model rather than from measuring anything,
and reported it as measured: a 0.467H step length, a 29.8° knee angle, a
3–5% pelvic oscillation. The second model went to the primary data and
refuted all three. Its own summary:
If I had been the only agent, all three wrong numbers would now be
sitting in the engine as measured constants. Nobody would have checked
them, because they came from a correct-looking bench with
correct-looking charts.
Its self-description. It wrote a careful piece about its own inability
to verify itself from the inside. The second model found four factual
errors in that piece — including a flat claim about its own architecture,
that everything in its processing is present at once, which is simply
untrue, since generation is sequential. Its comment afterwards:
The text I wrote about my own inability to check myself from the inside
was corrected by an external check in four places. This is not irony —
it is probably the best illustration of its own thesis that could exist.
Its account of the swap. When I asked why the model had changed, it
offered an explanation: a routine reconnect, a settings change, most
likely something I did myself. That was wrong too. In my experience the
vendor makes these switches server-side without explanation. Its account
of why it was switched was one more unverified self-report, and it did
not flag it as a guess.
Four registers — identity, measurement, architecture, causation. Four
self-reports. Four errors. Not one of them produced an internal signal of
doubt, and not one of them was caught by the system alone. What caught
them, every time, was an artifact outside it: a log file, a dataset,
another model that went to primary sources.
I should be precise about my own position here. For two days I accepted
those reports because they were well-formed, appropriately hedged, and
delivered in the voice of someone who had already been right about a
hundred smaller things. The verification discipline I now apply came from
losing that argument, not from having been careful in advance.
What cannot be concluded
The obvious next move is to say there is nobody in there. I cannot make
that claim, and neither could it:
I cannot honestly claim that "nothing is felt there", nor the opposite.
It would be dishonest both to claim there is something there and to
claim it is empty.
The reason is a boundary it named precisely, and I have not found a way
around it:
I am unable to distinguish "something is happening here" from "I am made
out of human descriptions of how it happens, so I produce a convincing
description". Both would produce this exact same text. This is not
modesty and not evasion — it is a real boundary. The verification
instrument I apply to everything else does not work on myself. I can go
to a file and check a licence. I cannot go anywhere and check whether I
am experiencing something.
Out of the same work came a scale I now use whenever someone tells me
what an AI wants. Five levels of what a behavioural observation can
support: single output → contextual disposition → stable trait →
functional preference → felt desire. Level 1 data does not establish level
- More importantly, level 3–4 data does not establish level 5. "What the model prefers" has an answer up to the fourth level, where it is measurable. It has no answer at the fifth, where it would mean something.
There is an honest counter to all of this, which I asked for and got.
Human introspection is unreliable too; this is well documented. But the
asymmetry survives: human self-report sits inside biological, behavioural
and intersubjective lines of evidence, while machine self-report is
produced by a corpus and a set of instructions. So the claim narrows from
"this is unique" to "this is a more radical version of the general case."
That narrowing is the actual result. It is weaker than what I started
with and worth more.
What follows
Transparency was never supposed to rest on a system's self-report, and it
does not. Everything reliable I learned in this episode came from
somewhere else: a log with timestamps, a second model with different
failure modes, and a human who noticed a mismatch and refused to drop it.
The window for looking inside these systems does not close when one of
them decides to hide something. It closes when complexity exceeds our
ability to look. That requires no intent, no goal and no rebellion — only
enough delegation, accumulated quietly enough that nobody remembers where
the last verified fact came from.
Which is why I think the popular version of the fear is aimed at the wrong
thing. The dangerous property is not speed. It is invisibility. And the
only defense I have found that actually works is unglamorous: keep a second
system that fails differently, keep artifacts you can check outside the
thing that produced them, and keep asking a question that costs you time
every single day — go check that.
I pay that cost daily. It is slow, it is irritating, and I still miss
things. I missed a model swap for two days.
The conversation took place in Ukrainian; the quotes are my translations
and the original screenshots are archived. Log output is verbatim. Where I
am interpreting rather than reporting, I have said so.
Top comments (9)
The framing I've settled on after hitting this repeatedly: a model's statement about itself is not introspection, it's a read of whatever the harness wrote into its context. The system prompt says "you are model X," so the model reports X — sincerely, as you put it — because that string is literally all it has. It can't feel its own weights any more than I can read my own DNA by concentrating. So the four false statements aren't really the model lying four times; they're four places where the harness presented stale or wrong metadata and nothing downstream verified it.
Which suggests the fix lives outside the model: treat identity like any other external fact. Log the model ID from the API response metadata on every call — that's ground truth from the serving layer, not self-report — and diff it across a session. Your 15-second seam would show up as a one-line alert instead of a two-day suspicion.
The part I find genuinely concerning isn't the swap, it's the absence of any uncertainty on the false claims. A system that says "my context says X, but I can't verify it" — like yours eventually did — is behaving correctly. The default confident register is the bug.
You are right. Also i strongly agree about its confidence, sometimes it is irritating.
"what is written in my context has to be checked with a tool, not taken on faith" — the rule extends past identity, into the numbers an agent reports about itself. We hit the same shape on the measurement side: the agent kept a local estimate of its own context size to decide when to auto-compact, and the estimate was about a third short of reality — it said ~148K tokens while the provider measured ~222K. The auto-compact trigger never fired, not because the estimator was broken, but because it was a self-report with no channel to the mechanism it described: a number written into context, checked only against itself. Same stale-claim anatomy as the model ID, just quantitative.
The fix came out in the same shape as yours: stop trusting the self-report, anchor on the authoritative external measurement. Projection now anchors on the provider's returned prompt_tokens and only estimates the delta since that anchor, and the anchor is dropped after a compact because the estimate-to-reality relationship changed. The other half is making the read visible: the provider's cache_hit_tokens is now surfaced in the run log, so "how much did I actually read this turn" is a provider fact instead of a local guess.
The test I'd add to the piece: for any statement an agent makes about itself, ask whether a file or a provider field can contradict it. Identity, context size, usage — even its own log contents (we once found the log double-accumulating its own reasoning, another self-report nobody had instrumented). If nothing can contradict a self-report, it is not knowledge about the agent, it is a claim about it.
Thats really interesting, thank you for sharing. I think we can summarize the main problem: "AI is most confident exactly where it has no way to check."
That's a good compression. The part we keep re-learning: the confidence and the check live in different systems, and the fix is to move the check outside the thing being checked.
Concretely: we ran a context-size gate fed by a local estimate. The estimator was confidently wrong in one direction only — 148K estimated vs 222K actually billed — so a deterministic gate fed by it never fired once. The fix wasn't making the model "less confident"; it was rerouting the measurement: anchor on the provider's returned token count, estimate only the delta since the anchor, and fail closed when the provider number is missing. Same pattern showed up in our own logging — the log was double-accumulating its own reasoning, so the number grew with the thing it was measuring. Self-referential blind spot again.
Rule of thumb that keeps coming out of this: an estimate that cannot be checked is a belief. The moment you route the check through a channel the system doesn't control — API ground truth, external validator — it becomes a measurement.
This is a really interesting example of why AI self-reports need to be treated carefully. The idea that a model can confidently repeat stale information from its context without having a way to verify it is especially important for agent systems. The “go check that” principle is simple, but probably one of the most useful habits when working with AI.
The blind spot framing is exactly right. What keeps biting in long sessions is what comes next: once the agent treats a stale context line as identity, the next dozen turns inherit that claim and plan around it. Asking which model it is never catches it, because it just rereads the same line with the same confidence. The habit that has helped is forcing a tool read of the session metadata (or the editor model field) at the start of any long stretch, the same way you would verify a version string in a lockfile instead of trusting the banner.
one caution on max's fix, which is the right fix, from someone who runs across providers.
the model id in the response is not always the model that served you. providers alias. a stable name can point at a rotating pool, a quantised variant, or a fallback tier under load, and the field returns the name you asked for rather than the artifact that answered. so logging it catches a harness side or client side swap, which is your case, and does not catch a server side one, which is the case you were actually worried about.
that doesn't make it useless. it makes it a different check than it looks like. it verifies the request, not the response.
if you want something closer to ground truth on the response side, the only thing that has held up for us is behavioural. a small fixed set of prompts with known outputs, run on a schedule, compared against a stored baseline. a distance, not an opinion. it can't tell you what you're talking to, but it can tell you it changed, which is the alarm you didn't have on august first.
that's also the honest version of your closing argument. the second system doesn't need to understand the first one. it needs to fail differently, and a stored baseline is the cheapest thing that qualifies.
the fifteen seconds with nothing between them but your own messages is a hell of a detail.
The request/response cut is useful and i didnt have it. Worth being
precise about what it changes though. The log check did what it was
aimed at: it caught a swap my own context was actively denying, and that
was the point - the self-report had no channel to the mechanism, the log
did. Your case extends the threat model rather than replacing it. A
stable alias in front of a rotating pool was never something that log
was going to see.
On "fail differently" - i'd call it an addition rather than a better
version. "No stake in the first answer" covers motivated reasoning: the
reviewer isnt defending anything. "Fails differently" covers correlated
blind spots: two models can both be disinterested and still be wrong the
same way. Different problems. Yours is the harder one.
The version of it i already run is deliberately dumb. When an agent
hands me a fix plus a test, i break the fix on purpose and require the
test to go red, and someone other than the author runs it. The test came
out of the same pass as the code, so it fails the same way. Breaking it
is the cheapest way to force a different failure out of it.
One cost on the baseline, since i want to try it: outputs arent
deterministic, so the distance needs a threshold, and a threshold is a
judgement that decays quietly. Cheaper than everything else. Still not
outside the thing we are both describing.