You upload your company handbook, or a folder of PDFs, or two years of notes. You ask it something the document plainly answers. It answers confidently, and it's wrong.
Then you do what everyone does. You rewrite the question. You add "only use the document provided". You try a better model. Sometimes it works, and you never find out why.
There's a way to find out why, and it takes about twenty minutes. Nothing to install, nothing to buy, no account.
First, what's actually happening
When you point an AI at your own documents, it doesn't read them the way you do. It can't — they're far too long to hold at once.
So the tool does something simpler than you'd expect:
- It cuts your documents into small pieces. Usually a few hundred characters each.
- When you ask a question, it picks the piece that looks most similar to your question.
- It shows the AI that piece, and asks it to answer.
The AI never sees your document. It sees one piece of it, chosen by a search you didn't configure and can't watch.

This is the actual situation. The scrap is what it answers from. The stack is everything it never saw.
Almost every wrong answer is decided at step 2, before the AI is involved at all. Which means arguing with the AI can't fix it.
The industry name for this arrangement is RAG. You don't need the term. You need to see step 2.
A worked example you can run
Here's a short leave policy. Forty lines of Python, no installs, no key — it does steps 1 and 2, then stops and shows you what the AI would have been given.
Save it as check.py and run python check.py:
# Shows what your AI actually gets handed when you ask about your documents.
# Python 3. Nothing to install, no API key, no account.
CHUNK_SIZE = 350 # <-- change this number and run it again
DOCUMENT = """
Annual leave
Every full-time employee gets 24 days of annual leave a year, plus public
holidays. Leave is counted from 1 January.
Requests go through the HR portal. Your manager approves them. If your manager
is away, their deputy approves instead.
How much notice you need to give depends on how long you are taking off. For a
single day, give two working days of notice. For anything up to a week, give
two weeks of notice. For longer than a week, give one month of notice.
Carrying leave over
You can carry a maximum of 5 unused days into the next year. Anything above 5
days is lost on 31 December. Carried days must be used by 31 March.
"""
QUESTION = "how much notice do I need to book a week off?"
def chunk(text, size):
words, chunks, current = text.split(), [], ""
for w in words:
if len(current) + len(w) + 1 > size:
chunks.append(current.strip())
current = ""
current += w + " "
if current.strip():
chunks.append(current.strip())
return chunks
def score(text, question):
stop = {"how", "much", "do", "i", "need", "to", "a", "the", "for", "of", "is"}
q = {w.strip("?.,").lower() for w in question.split()} - stop
c = {w.strip("?.,").lower() for w in text.split()}
return len(q & c)
chunks = chunk(DOCUMENT, CHUNK_SIZE)
ranked = sorted(range(len(chunks)), key=lambda i: score(chunks[i], QUESTION), reverse=True)
print(f"chunk size {CHUNK_SIZE} -> {len(chunks)} pieces")
print(f"question: {QUESTION}\n")
for rank, i in enumerate(ranked, 1):
print(f" #{rank} piece {i} score {score(chunks[i], QUESTION)}")
print("\nThis, and only this, is what the AI gets:\n")
print(f" {chunks[ranked[0]]}\n")
print("Read it. Could you answer the question from that alone?")
The document says, in plain English, that a week off needs two weeks of notice. Here's what comes back:
$ python check.py
chunk size 350 -> 2 pieces
question: how much notice do I need to book a week off?
#1 piece 0 score 2
#2 piece 1 score 2
This, and only this, is what the AI gets:
Annual leave Every full-time employee gets 24 days of annual leave a year, plus
public holidays. Leave is counted from 1 January. Requests go through the HR
portal. Your manager approves them. If your manager is away, their deputy
approves instead. How much notice you need to give depends on how long you are
taking off. For a single day, give two
Read it. Could you answer the question from that alone?
Look at the last five words.
For a single day, give two — and then the text stops. The sentence was cut in half.
An AI handed that will answer "two days", fluently and with no hedging. It isn't hallucinating. It's answering correctly from the only evidence it was given. The evidence was just the wrong half of a sentence.
And notice the scores: both pieces scored 2. It was a tie. The wrong one won because it happened to come first.
Now change one number
Set CHUNK_SIZE = 300. Same document, same question, same code:
$ python check.py
chunk size 300 -> 3 pieces
question: how much notice do I need to book a week off?
#1 piece 1 score 3
#2 piece 0 score 1
#3 piece 2 score 0
This, and only this, is what the AI gets:
long you are taking off. For a single day, give two working days of notice. For
anything up to a week, give two weeks of notice. For longer than a week,
give one month of notice. Carrying leave over You can carry a maximum of 5
unused days into the next year. Anything above 5 days is lost on 31
Read it. Could you answer the question from that alone?
Correct. Because the piece boundary landed somewhere else.
Here is every size I tried:
chunk size pieces answer correct?
150 5 yes
200 4 NO
250 3 yes
300 3 yes
350 2 NO
400 2 NO
450 2 yes
500 2 yes
That isn't a curve you can tune your way along. It's arbitrary. Right, wrong, right, right, wrong, wrong, right, right.
A number you never chose, and probably never saw, decided whether your answer was true. Not the model. Not your prompt.
The test
That's the whole point of the exercise, and it generalises. Whatever tool you're using, before you touch the prompt:
- Find what the tool actually retrieved. Most of them will show you — it's called sources, citations, references, or context. Open it.
- Read that text yourself. Ignore the answer entirely.
- Ask one question: could I have answered correctly from this text alone?
That single question splits every wrong answer into two completely different problems:
No, I couldn't have — then the AI never stood a chance. The retrieval is broken. Rewriting your prompt is wasted effort, and this is where most people spend their afternoon.
Yes, I could have — the answer was right there and the model went past it. A genuinely different problem, with a different fix.
Twenty minutes. No framework, no monitoring platform, no spend.
If it's the first one
Things that actually help, roughly in order of effort:
Ask for more pieces at once. Most tools fetch one. Three or five means a cut sentence gets rescued by its neighbour.
Cut on meaning, not on character count. Split at paragraphs and headings. A rule that ends a piece mid-sentence will eventually end one mid-answer.
Keep the heading with the text. A piece that says "give two weeks of notice" without "Annual leave" above it can't be matched to a question about leave.
Don't rely on similarity alone. Plain word matching catches the things similarity misses — product codes, error numbers, names. Using both together is now the ordinary recommendation, not an advanced one. NovaKit make the case well, and kapa.ai list the other common mistakes.
Those guides are good and I'd read both. What none of them gives you is the part above: which of the five things is yours. They hand you a list of suspects. The read-it-yourself test hands you a defendant.
The thing nobody tells beginners
Every article on this ends at "improve your retrieval". Fair enough. But there's a step past it, and it changes how you think about the problem.
Better chunking makes wrong answers rarer. It never makes them impossible. You're tuning a dial and hoping.
On a system I built — a multilingual assistant answering questions across about 75,000 lines of source — we stopped tuning the dial. Every claim in every answer had to resolve to a real, checkable location in the source material. If a citation didn't resolve, the answer didn't ship. Not flagged. Not scored. Withheld.
The goal isn't an AI that's usually right. It's a system where being wrong is structurally impossible, because a claim it can't back up never reaches you.

Checked against the source, or it doesn't get handed over. Withheld, not flagged.
That's a design decision, not a setting, and it's the difference between a demo and something people can rely on at work. You don't need it on day one. But it's worth knowing the ceiling exists, so you stop expecting prompt changes to get you there.
Start here
Next time your AI gets your own documents wrong, don't rewrite the question.
Look at what it was given, and ask whether you could have answered from it.
Most of the time you couldn't have — and the machine you were about to blame was the only part doing its job.
Check #2 above is one of eight. I've put the rest on a single printable page — the things worth thirty seconds before you act on what an AI told you. Free, no signup: Before you trust an AI answer →
I build this kind of thing for a living, and I give away the tools that check it — all of them here, free and readable. What I've built · work with me.
Run the script on your own document and something surprising falls out? Tell me — I collect these.
Read next: What is Jev? The AI trained to say “I’m only 60% sure”
Originally published at singhlabs.dev.
Top comments (2)
The tied scores in the 350-character example make the failure easy to inspect without blaming an invisible model choice. I would qualify the later claim about structural impossibility, though: a citation resolving to a real source location does not establish that the source entails the generated claim.
A useful test would pair a valid location with an incorrect interpretation, such as citing the one-day notice rule while answering the week-long leave question. That separates citation resolution from claim support. Withholding unresolved citations is a valuable gate, but the supported-claim check still needs to catch scope, exceptions and contradictory passages before the answer ships.
The forty-line script isolating the retrieved slice before the LLM even sees it is the clearest demonstration of this failure mode I have seen.
The other place where naive character chunking breaks down is structured markdown tables and key-value lists. If a row header lands in chunk one and the numbers land in chunk two, semantic search fetches one or the other, but never both together. Moving to parent-child retrieval solved that for us, embedding small paragraph-level chunks for search precision, but handing the full parent section to the model so sentences and tables never arrive amputated.