DEV Community

Dimitris Kyrkos
Dimitris Kyrkos

Posted on

Your model doesn't need more training. It needs a better search index.

Intro

Every team that plugs an LLM into its business hits the same moment. Someone asks the model about an internal pricing rule, a product SKU, or a clause in the standard contract, and it answers confidently and wrongly.

The reflex that follows is almost universal: "The model doesn't know our business. Let's fine-tune it."

So the team spends weeks exporting tickets and docs, cleaning them, formatting them into JSONL, and paying for training runs. The new model sounds more like the company. It uses the right jargon. And it still makes things up.

That's not bad luck. It's a category error.

What fine-tuning is actually good at

Fine-tuning adjusts a model's weights so it behaves differently. That's powerful when the thing you want to change is behavior:

  • A specific tone or writing style (support replies that sound like your brand, not like a chatbot)
  • A strict output format (always return this JSON schema, always produce this report layout)
  • A narrow, repetitive task where a smaller tuned model can replace a bigger general one

Notice what all of those have in common. They're about how the model answers, not what facts it knows.

When the goal is "the model should know our refund policy, our product catalog, and last quarter's changes," you're asking weights to act like a database. They're a bad one.

Failure mode 1: it still hallucinates, just in your accent

A fine-tuned model doesn't gain a reliable sense of what it knows and what it doesn't. Training on your documents shifts probabilities toward your vocabulary, but when a question lands in a gap, the model does what it always does: it produces the most plausible-sounding continuation.

Illustrative scenario: a support assistant tuned on two years of tickets is asked about a warranty extension introduced last month. It has never seen it. It answers anyway, in perfect company voice, with the terms of the old warranty. The answer is more convincing than a generic model's would have been, which makes it more dangerous, not less.

Fluency in your domain is not the same as accuracy about your domain.

Failure mode 2: every fact change is a training run

Business knowledge isn't static. Prices change, policies get revised, products get deprecated, a regulation lands and three documents get rewritten.

If that knowledge lives in the weights, every update means rebuilding the dataset, retraining, re-evaluating, and redeploying. In practice, teams don't do that weekly. So the model drifts out of date, quietly, and nobody knows exactly which facts are stale.

Compare that with a retrieval setup: a document changes, you re-index it, and the next query sees the new version. The update cycle is minutes, not a project.

Failure mode 3: you can't show your work

When an answer comes from retrieval, you can point to the exact chunk of the exact document it was grounded in. You can show it to the user. You can log it. When someone disputes an answer, you can check whether the source was wrong or the model misread it.

When an answer comes from fine-tuned weights, there is no source. The knowledge is smeared across billions of parameters. You can't cite it, you can't audit it, and you can't tell a compliance team where a claim came from.

For anything customer-facing, regulated, or contractual, that alone should end the discussion.

What to build first instead

Before training custom weights, invest in the boring part: retrieval.

user question
   -> query rewriting / expansion
   -> hybrid search (keyword + embeddings) over a clean, deduplicated index
   -> reranking
   -> top-k chunks + source metadata into the prompt
   -> answer with citations
Enter fullscreen mode Exit fullscreen mode

Most of the quality comes from things that have nothing to do with the model:

  • Clean source documents (no three conflicting versions of the same policy)
  • Sensible chunking that keeps related context together
  • Metadata (dates, owners, product, region) you can filter on
  • Hybrid search, because embeddings alone miss exact identifiers like SKUs and error codes
  • An evaluation set of real questions with known correct answers, so you can measure whether changes help

None of this is glamorous. All of it pays off more than a training run for knowledge-heavy use cases.

When fine-tuning does make sense

This isn't "never fine-tune." It's "fine-tune for the right reason, and usually later":

  • Retrieval is solid, answers are grounded, but the output format or tone keeps drifting. Tune for behavior.
  • You need a smaller, cheaper model to handle a narrow, high-volume task a big model already does well. Tune for cost.
  • The model struggles to use retrieved context well in your domain (for example, dense technical or legal text). Tune on examples of reading and citing context, not on the facts themselves.

In all three cases, facts still come from the index. The weights handle behavior.

The reframe

Fine-tuning teaches a model how to talk. Retrieval tells it what's true right now. Most "the model doesn't understand our business" problems are the second kind wearing the costume of the first.

What pushed your team toward fine-tuning (or away from it), and did it actually fix the knowledge problem you were trying to solve?


Top comments (23)

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

The line I want to put numbers against is "hybrid search, because embeddings alone miss exact identifiers like SKUs and error codes." We measured that on our own deployed store yesterday. It held. Then it broke the next step of your advice, in a way that surprised me.

Ten answerable questions, gold document known, document level recall over our own engineering notes:

retrieval arm R@10 (the depth we serve) R@100
dense only 0 of 10 3 of 10
BM25 only 3 of 10 7 of 10
hybrid RRF, as we had it configured 0 of 10 4 of 10

Our recovery runbook ranks first under BM25 and appears nowhere in dense's top 100. Our notes are full of filenames, flags and error strings, which is the shape you name.

The surprise is the third row. Our hybrid beat dense and still lost to plain BM25, for two reasons we own. An admission gate decided which queries were specific enough to deserve the lexical arm and let through 7 of 14, and among the ones it rejected was the single query where BM25 ranked the gold document first, so that one fell back to dense and missed. Then, where BM25 was admitted, RRF at K=60 diluted it. Sparse rank 6 landed at hybrid 11, sparse 11 at 16, and sparse 33 at a miss. Fusing a strong lexical list with a near random dense list spends rank mass on noise.

Underneath both, our full text sidecar was stale, so production had been serving dense only while the config said hybrid. Your clean index point arrives here as an outage, and nothing in our monitoring saw it, because ten plausible chunks look the same to a dashboard either way.

So one line I would add to your build list. The eval set is what tells you which arm of your hybrid is actually running. Ours told us, and it took ten questions with known answers.

Caveats. Ten usable tasks out of fourteen, private corpus, document level, single run. No LLM judge anywhere in it, so those counts are deterministic. And the public check runs the other way. On SciFact our whole document dense arm and a standard chunked arm came out level at nDCG@10 0.7014 against 0.7016, with a BM25 control at 0.6644 matching the published figure, which is how we know the harness is honest. On the public set the choice barely moved anything. On our own corpus it decided the outcome.

Collapse
 
nomad-link-id profile image
Igor Eduardo •

This is the category error I keep seeing too: teams fine-tune because the model "doesn't know our business," when the real gap is an index that can't surface the right clause, SKU, or error string today.

Weights are a bad database. Fluency in your accent after a training run is not the same as an auditable evidence path.

Tom's receipts on hybrid are the part I'd underline for anyone shipping this: "hybrid on" in config is not hybrid in production. If an admission gate drops the lexical arm on the exact query where BM25 would have won, or RRF dilutes a strong sparse list with near-noise dense ranks, the blended score hides a dead arm. The fix isn't a bigger model — it's a tiny golden set with known documents, scored per query class (identifier / exact string vs semantic paraphrase), so you can see which arm actually ran.

Retrieval can suggest. It shouldn't silently decide. And if you can't point at the chunk, you don't have a knowledge system — you have a plausible voice.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

That breakdown of how hybrid search actually breaks in production is spot on, especially the way reciprocal rank fusion can dilute a strong sparse signal with dense noise. It is too easy to check the hybrid box in a configuration file and completely miss that one of your retrieval arms is effectively dead for specific query classes. Your suggestion to use a targeted golden set to audit which arm is actually winning per query type is a much more practical way to debug retrieval than just staring at aggregate evaluation scores.

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

"Weights are a bad database" is the compression of the whole argument, and I am going to steal it.

Two things back, and the second one argues with your fix instead of agreeing with it.

On "if you can't point at the chunk, you don't have a knowledge system." I would have said that this morning. Today I found the state below that one, and it is worse.

We push short notes to an agent bound to the action it is about to take. I shipped the same bug twice in one session, then went looking for the note we must be missing. The note existed. It was selected for exactly that action. I could point at the chunk. And what got delivered was 675 characters of 3,751, because the channel caps each item, so the sentence I needed sat past the cut. Corpus-wide that is 98,715 characters which are stored, correctly matched, and never delivered.

So pointing at the chunk is necessary and short of sufficient. There is a third state between missing and arriving: selected, pointed at, and truncated, which is indistinguishable from arriving at the point of use. Retrieval decided nothing silently there. Delivery did.

One honest correction to my own claim, made a few hours ago. I had been saying that third state "behaves exactly like missing," and I went and tested it properly rather than asserting it: reinstate the cut text, rerun the same blind judge, 65 cut pairs. Restoring it changed the judged action 9 times and lost it 5 times, p=0.42, and it did not beat a length-matched tail of unrelated text. So the 98,715 is a real character count and the harm is not yet demonstrated. Underpowered rather than refuted, but I am not entitled to the stronger sentence.

Now the argument. Your fix is a tiny golden set with known documents, scored per query class, and the per-query-class half is right, because that is the same shape as everything else here. The word "tiny" is where I would push.

We ran six frontier models over the same tool-calling set at two sizes. Same models, same harness, same day.

items per model model pairs that separate
220 0 of 15
800 8 of 15

At 220 the honest reading of that table was "these are all the same". At 800, one of them turns out to score 61% on one category against 87 and 88 on the other two. The small set gave the opposite answer rather than a weaker one, and carried no sign that it was underpowered.

A golden set small enough to maintain is also small enough to report "both arms look fine" while an arm is dead. So I would keep your stratification and add a stopping rule to it: per query class, how many known-answer queries before a dead arm would actually show up. For us the answer came out well above ten, and ten was what we had.

Our own version of the thing you are pointing at, in case the number is useful: we resolved every top-1 retrieved document against the file that actually held the answer, and it was the right document 0 times out of 14. The blended score looked healthy the whole time.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

This is the kind of receipts I wish more people posted. The RRF-at-K=60 dilution point is especially good, because it's the exact failure mode people don't expect: fusing a strong list with a weak one can be worse than just using the strong list alone. The admission gate story is painful too, since the one query where BM25 would have won got routed away from it. Classic case of the meta-logic around retrieval being the actual bug, not the retrievers themselves.

And your last point is the one I want to steal. "The eval set tells you which arm of your hybrid is actually running" is a better argument for evals than anything I wrote. A stale full-text sidecar silently degrading you to dense-only is exactly the kind of outage that dashboards will never catch, because the output shape is unchanged. Ten known-answer questions caught what monitoring couldn't. That's the whole pitch for goldens right there.

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

"The eval set tells you which arm of your hybrid is actually running" is the part I would keep too, and I spent today finding the floor underneath it, which turns out to be lower than I thought.

The goldens tell you the retriever is alive. Nothing was telling me the goldens were alive.

We run 83 guards in this repo, each declaring its own negative control: a fixture carrying the exact defect the guard exists to catch, which the guard must refuse. There is a harness that runs every one of them and reports which guards can actually fail. Tonight it read 52 proven, 0 broken, 31 unproven. It read 50 proven and 2 broken when I started, and the three defects it found are all the same shape as your stale sidecar: the check ran, the check was green, and the green meant nothing.

One control set its instance fixture but not its volume fixture, so half the check still queried live AWS. It could never be shown passing on clean input, and the guard had been filed as broken for weeks. The guard was fine.

One control used a relative path. It worked from the repo root and crashed from anywhere else. The harness reported the crash, somebody read that as the guard failing, and again the guard was fine. A control that only works from one directory is reporting on your shell.

The third one was mine, shipped ninety minutes earlier. Its fixtures were scored against a corpus derived from recent commits, so the corpus moved while I worked. Run directly: the bad fixture scored 82% and was refused. Run from the harness eleven minutes later: 43%, accepted. Same fixture, same code, opposite verdict. I had written a weather report and called it a control.

So the version of your rule I am taking away is one turn further in. A golden set catches a silently degraded retriever. A negative control catches a silently degraded golden set. And the thing I did not have until today was anything that runs all the negative controls and tells me which of them have quietly stopped being able to fail.

The cheap part is that "unproven" is a real answer. Thirty-one of ours have no declared control and the harness says so plainly rather than counting them as passes. A number that admits what it has not checked is worth more than a green one that does not.

Thread Thread
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

The progression from golden sets to negative controls is a massive eye-opener, especially the realization that your evaluations themselves can quietly degrade and give false greens. Having a harness that actively tests if your guards can still fail, and being comfortable with an unproven status rather than a fake passing one, is next-level engineering for these systems. It highlights how the hardest part of building with language models is not the model itself, but building the scaffolding that proves your testing pipeline is actually telling you the truth.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

Thank you. One thing so it reads at its real size: when I wrote that, 31 of the 83 guards had no declared control at all, so for over a third of them the honest answer was unproven. Those guards are exactly as good as before. The harness stopped counting them as passes, which is a smaller change than it sounds and the one I would recommend first, because it costs nothing and changes what the green line means.

Thread Thread
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Exactly. “Unproven” is useful because it separates “the guard passed” from “we have evidence this guard can catch anything.” That distinction sounds small until you have enough checks that nobody remembers which ones were ever actually exercised. The same principle applies to evals too: a green result is only meaningful if the test itself has demonstrated that it can go red.

Thread Thread
 
spandaworks profile image
Tommie •

Agreed, and it cuts the other way too. Today one of our guards refused two whole benchmark runs because it treated every NameError as our harness's fault, when one was a model typo, total instead of total2. It had been shown it could go red; nobody had checked it went red for the right reason. The fix was to make it ask whether the missing name was something the prompt provided, with a test for both cases.

Collapse
 
dannwaneri profile image
Daniel Nwaneri •

Dimitris, I never got pushed toward fine-tuning. I built the retrieval stack first. Hybrid search, BM25 plus semantic, then a cross-encoder rerank stage on top.

Cloudflare highlighted that setup officially. Failure mode 3 is the one most people miss. Once you can show the exact source chunk behind an answer, a compliance dispute stops being a guessing game.

This post is timely too. Yesterday I spent 50 minutes on this video about building an LLM from scratch: youtu.be/YmLp8qe87A0?si=e0TyKZnyMs...

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Congrats on the Cloudflare feature, Daniel, that is a huge win. Starting with the hybrid and reranking stack is definitely the way to go, and you are so right about compliance. Having a clear paper trail to the exact source chunk saves so much debugging and legal pain down the road. Thanks for sharing that video link too, I will definitely have to check it out.

Collapse
 
constant_itis profile image
Constant Itis •

Meanwhile, sitting neglected in the corner:

function searchTheDamnDatabase(query) {
validate(query)
searchKnownSources()
enforceFilters()
checkVersion()
dedupe()
returnEvidenceWithProvenance()
}

😭

The recurring mistake is treating an LLM like the runtime environment instead of like a probabilistic reasoning/interface layer sitting on top of software.

If the process is knowable, testable, repeatable, and can be encoded:

encode it.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Ha, yeah, the number of times I've seen a team reach for fine-tuning when what they actually needed was a WHERE clause and a version column is genuinely depressing. The LLM-as-runtime confusion is the root of so much of this. It's a great interface and a mediocre database, and people keep asking it to be the second thing because the first thing is so impressive.

Your rule is the right one. If it's knowable, testable, and repeatable, encode it, and let the model do the fuzzy part on top. The model should be the last mile, not the whole road.

Collapse
 
ikrame-ih profile image
Ikrame Ibn Hayoun •

In one of my projects I mix plain SQL filtering with fuzzy matching and embeddings, because when money is involved an exact detail matters as much as semantic similarity. The line I keep defending: retrieval can suggest, but it shouldn't get to decide

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Yeah, that hybrid is exactly the right instinct when money is on the line. SQL for the parts that have to be exact (customer ID, account, currency, effective date), fuzzy and embeddings for the parts where the user's phrasing won't match your schema, and then let the model reason on top of a result set your code already trusts. The embeddings are there to help you find the row, not to be the row.

And your last line is the one I want to steal outright. "Retrieval can suggest, it shouldn't get to decide" is a cleaner version of what I was circling around in the post. The moment a retrieval score gets to determine which account gets refunded or which price applies, you have promoted a similarity metric into a business rule, and similarity metrics were not designed to carry that weight. Suggest, rank, surface candidates, fine. Decide, no.

Collapse
 
glenallen profile image
Glen Allen •

The gap between a retrieval architecture on paper and the retrieval system actually serving queries is an important failure mode here. A stack can be configured for hybrid search, reranking, and fresh indexing, while a stale index or routing decision quietly means production is using something else. I’d treat the retrieval pipeline itself as part of the evaluation target: not just “did the gold document appear?” but “which retrieval path produced it, what index version was queried, and did reranking actually improve the final position?” That kind of instrumentation could make retrieval failures much easier to distinguish from model failures, especially when the answer still looks plausible.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

That is a spot-on point. It is so easy to blame the LLM for a bad answer when the real culprit is a silent failure in the retrieval pipeline, like a stale index or a bad routing decision. Treating the retrieval path as an evaluation target and logging things like index versions and reranker performance is crucial. Without that level of telemetry, you are basically flying blind and hoping the right context actually made it into the prompt.

Collapse
 
brianainews profile image
Brian · AI News •

The most useful distinction here is that retrieval keeps facts updateable while fine tuning changes behavior. I would add a freshness check at retrieval time so stale indexes fail loudly instead of producing a perfectly cited wrong answer.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Agreed, and honestly that should have been on my build list. A freshness check plus a hard fail (or at least a visible degradation signal) is the difference between "we served a stale answer for three weeks" and "we noticed at 9:04am." The perfectly cited wrong answer is arguably the worst failure mode in the whole stack, because the citation makes people trust it more, not less.

Simple version: every chunk carries a last-indexed timestamp and a source-last-modified timestamp, and if the delta blows past a threshold, the retriever surfaces it instead of quietly serving it. Cheap to add, catches a whole class of silent drift.

Collapse
 
suraj09 profile image
Suraj Suradkar •

The distinction between model capability and retrieval quality is really important.

I think there's a similar issue with project context: having more information available doesn't necessarily mean the agent has the right context for the decision it's making. Stale or conflicting project history can be just as problematic as missing context.

The metadata and source-traceability point feels especially important when context needs to stay useful over time.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Yeah, the parallel holds really well. "More context available" and "the right context for this decision" are two different things, and they get conflated constantly. A sprawling project history that includes three superseded decisions, two abandoned approaches and one current plan is arguably worse than a smaller context that only carries the live state, because the agent has no principled way to know which layer is authoritative. Volume without provenance is just noise with extra tokens.

The metadata point is where I think this really lives. Dates, owners, status, supersedes-relationships, whatever your domain equivalent is. Without those, "retrieved" and "true right now" quietly drift apart, and you get the same failure mode as fine-tuned weights: a confident answer with no way to tell how stale the source was. Traceability is what lets you separate "the retriever failed" from "the retriever did its job and the source was out of date," and those need very different fixes.

Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more