DEV Community

Dimitris Kyrkos
Dimitris Kyrkos

Posted on

Half the AI agents in production are if-statements with a GPU bill

Real-world examples like regex-replacing prompts

There's a new kind of technical debt, and it doesn't come from cutting corners. It comes from reaching for the most impressive tool in the room.

Call it resume-driven AI engineering: picking an agent framework, a vector database, or a multi-model orchestration layer because it looks great on a CV, not because the problem needs it. The result works in the demo. It's also slower, more expensive, harder to debug, and nondeterministic in places where it didn't need to be.

The demo vs. the pager

In a tutorial, complexity is free. You spin up an agent, wire in a vector store, and watch it do something clever with ten sample documents.

In production, every moving part has a cost: latency, token spend, a new failure mode, a new thing someone has to understand at 3 a.m. The question isn't "can an LLM do this?" (it usually can). It's "is an LLM the simplest thing that does this reliably?"

Here are three places where the answer is often no. The scenarios are illustrative, but if you've been around AI projects for a while, they'll look familiar.

Failure mode 1: a model call where a regex would do

A team needs to pull invoice numbers out of incoming emails. Invoice numbers follow a fixed format: INV- plus eight digits. They send every email to an LLM with a prompt asking it to extract the number.

It works 98% of the time. The other 2%, the model "helpfully" reformats the number, or picks up a purchase order number instead. Each call costs money and adds a few hundred milliseconds.

import re

INVOICE_RE = re.compile(r"\bINV-\d{8}\b")

def extract_invoice_ids(text: str) -> list[str]:
    return INVOICE_RE.findall(text)
Enter fullscreen mode Exit fullscreen mode

Deterministic, testable, effectively free, and it runs in microseconds. Keep the model for the messy cases the pattern can't handle, and route to it only when the regex finds nothing.

Failure mode 2: vector search where SQL would do

"Show me all orders from customer 4417 in the last 30 days that are still unpaid."

That's not a semantic question. It's a filter. Yet it's common to see this kind of query embedded, pushed through a vector store, and answered by an LLM summarizing the top-k chunks, which may or may not include every matching order.

SELECT id, total, created_at
FROM orders
WHERE customer_id = 4417
  AND status = 'unpaid'
  AND created_at >= NOW() - INTERVAL '30 days';
Enter fullscreen mode Exit fullscreen mode

Exact, complete, indexed, auditable. Vector search is great when you're matching meaning ("tickets similar to this complaint"). It's the wrong tool when you're matching facts.

Failure mode 3: an autonomous agent where a decision tree would do

A support workflow: if the customer is on the enterprise plan and the issue is billing, route to account management; if it's a bug, open a ticket; otherwise, send the FAQ link.

That's four branches. Someone builds it as an autonomous agent with tool access, a planning loop, and a memory store. Now the routing is probabilistic, occasionally loops, and nobody can explain why ticket #8812 went to the wrong team.

def route(customer, issue):
    if customer.plan == "enterprise" and issue.type == "billing":
        return "account_management"
    if issue.type == "bug":
        return "open_ticket"
    return "send_faq"
Enter fullscreen mode Exit fullscreen mode

If you can draw the logic on a whiteboard, you probably don't need an agent to rediscover it every request.

The reframe

A lot of "AI systems" are really ordinary software with an LLM bolted onto a step that didn't need one. The model isn't the problem. Using it as the default instead of the exception is.

A practical decision checklist

Before adding a framework, a model call, or an agent, ask:

  • Is the input structured or the output fixed-format? Start with parsing, regex, or schema validation.
  • Is the question about facts or about meaning? Facts go to SQL. Meaning can go to embeddings.
  • Can the logic be enumerated? If yes, write the branches. Agents are for open-ended tasks where you genuinely can't.
  • What happens when it's wrong? If the answer is "silent bad data," you want determinism.
  • Who maintains this in a year? Every framework is a dependency someone has to upgrade, understand, and debug.

None of this means "never use AI." Use it where ambiguity actually lives: unstructured text, fuzzy matching, generation. Just make it the tool you reach for on purpose, not by reflex.

The best engineers aren't the ones with the most complex stack. They're the ones whose systems are still simple enough to understand when something breaks.

What's the most over-engineered AI setup you've seen (or built) that could've been replaced with something boring?

Top comments (23)

Collapse
 
ingosteinke profile image
Ingo Steinke, web developer •

I don't see the traditional exact coding as boring at all. Take regular expressions for example: regular, compact, but highly complex.

Thanks for the specific examples and code snippets! Your overall approach resonates with the principle of least AI and several proven UNIX philosophies. Prefer one simple tool that does one thing well, don't grant unnecessary permissions, don't overengineer, so to say.

The most overengineered setups I see are all those harnessing demos right now. People write a wishlist in their agents file, then they let AI modify the code hopefully according to the written requirements, and run tests and linters after each iteration. They still need a "human in the loop" to review and fix and tighten the ruleset. Before that scales, they could probably have written everything by themselves in the same time and with better security.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

True, regex is definitely an art form in itself. That agentic code-generation loop is a classic case of spending ten hours automating a task that takes ten minutes to write. It basically replaces actual engineering with a massive QA and debugging cycle for mediocre code. Sometimes just sitting down and writing the code is still the fastest, most secure path to production.

Collapse
 
eternaclarity profile image
Jesse Gamble •

Asking what happens when it's wrong is the question that separates the checklist from the hype. For teams inheriting an over-built agent, a cheap first cut is logging the agent's intermediate decisions for a week, then hard-coding the branches that never vary. Silent failures are what make the 98% invoice case so expensive.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

The logging idea is especially useful for inherited systems. You can learn a lot by watching what the agent actually does before ripping anything out.

I’d probably be a little careful about hard-coding branches just because they stayed the same for a week, though. That could be a useful signal, not necessarily proof that the branch is stable.

Collapse
 
syntaxwanderer_26 profile image
Taras Hanych •

Strong agree on "is an LLM the simplest thing that does this reliably?". I'd extend it to the coding side too: a lot of teams now ask an agent to remember architecture rules from a prompt, when a plain deterministic check would enforce them for free and never have a 2% failure rate. Same principle as your regex example: keep the model for the fuzzy part and turn everything that can be a rule into code. The boring solution is usually the one that survives the 3 a.m. page.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Yeah, the coding side is a good extension of this. Especially when the rule is something like "this package can't import that package", there isn't much value in asking a model to remember it when a linter or CI check can just reject it.

I think the interesting cases are where the rule is partly fuzzy. That's probably where the model earns its keep, rather than making it responsible for enforcing rules that can be expressed directly in code.

Collapse
 
nomad-link-id profile image
Igor Eduardo •

Failure mode 2 is the one I keep seeing in "knowledge" bots: a filter query gets embedded, top-k'd, and summarized — then someone is surprised the unpaid-orders list is incomplete.

That isn't a retrieval quality problem. It's a category error. Customer 4417 + last 30 days + unpaid is a closed-world fact query. SQL (or any indexed filter) is the evidence path; a vector hit list is a probabilistic shortlist that was never asked to be complete.

Same split on identifiers and fixed formats: regex/schema first, model only on the residue. Keep the LLM where ambiguity actually lives — paraphrase, messy prose, open-ended planning — not where a wrong digit silently ships.

Practical check I'd add to your checklist: for each production question class, mark exact / filter / semantic. If the class is exact or filter and the path still goes through embeddings + an agent loop, the GPU bill is paying for nondeterminism you didn't need.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Yeah, I like the exact/filter/semantic split. It makes these decisions a lot easier to reason about than starting with "which retrieval stack should we use?"

The interesting bit is that sometimes the vector path gets introduced so early that nobody stops to ask whether completeness is even a requirement. Once you frame it that way, the tradeoff gets pretty obvious.

Collapse
 
brianainews profile image
Brian · AI News •

The invoice regex case is the one I keep seeing get skipped. Teams treat 98 percent extraction as good enough and never measure the 2 percent that quietly rewrites the number. Routing to the model only when the pattern misses is the right split, but I would also log those misses as a labeled set. After a month you can see whether the leftover cases are actually messy or just a second pattern you never wrote.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Yeah, logging the misses changes the fallback from "AI handles the weird stuff" into something you can actually learn from. After a while you might discover that half the supposedly messy cases are just another predictable format.

That also gives you a nice feedback loop for deciding whether the model is still earning its place in the pipeline.

Collapse
 
makeyouragent profile image
MakeYourAgent •

The unpaid-orders example is the failure mode I see most on knowledge and support bots.

Customer 4417, last 30 days, unpaid is not a semantic neighborhood. It is a filter. Embed it, top-k it, summarize it, and you get a confident incomplete list. That is not weak retrieval. That is the wrong tool.

Same with the invoice regex. If the shape is fixed, parse first and call the model only on misses, and log those misses. The 2% that quietly rewrites the number is worse than a clean reject.

For internal SOP and help-center bots I keep identifiers, statuses, and date windows as code or SQL, and reserve the model for wording once the rows are already correct. Demo complexity is cheap. Pager complexity is not.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

Yeah, the "confident incomplete list" part is what makes the vector approach particularly awkward here. A bad semantic match at least feels like a retrieval problem. Missing records from a query that should be exhaustive is a different kind of failure.

I like the idea of keeping the model on the wording side once the actual rows are correct. It gives you a much cleaner boundary for testing too.

Collapse
 
hannune profile image
Tae Kim •

The quiet reformatting failure never shows up in the model's confidence score, which is what bit us. We had entity resolution pipelines where an invoice number would come back with a transposed digit or a missing hyphen, the match looked fine, then two reconciliation cycles later a join that should have been deterministic was not. Pulling fixed-format extraction out completely and only calling the model on genuinely ambiguous inputs cut that failure class to zero. Should have done it earlier but the 98 percent accuracy during eval had looked good enough.

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

That's a nasty failure mode because the output still looks reasonable to a human. A clean rejection is much easier to notice than a valid-looking identifier that's subtly wrong.

The reconciliation example is a good reminder that 98% accuracy isn't necessarily the right metric for fixed-format data. Sometimes the important number is how often you produce a wrong value instead of saying "I don't know."

Collapse
 
mateo_ruiz_6992b1fce47843 profile image
Mateo Ruiz •

The “make the model the exception, not the default” principle is probably the most important takeaway here. The interesting architectural question isn't whether an LLM can perform a task, but whether introducing probabilistic behavior actually buys enough value to justify the additional failure surface.

I’d add one more dimension: where should uncertainty live? At IT Path Solutions, when designing production AI workflows, we try to keep deterministic boundaries around things like authorization, state transitions, validation, and transactional operations, while letting the model handle the genuinely ambiguous parts. That separation makes failures much easier to isolate.

There’s also a subtle benefit to this approach: simpler components give you better observability. If an invoice parser, SQL query, or routing rule fails, you can usually reproduce the exact input and reason about the failure. With an autonomous loop, the same bug can depend on model output, tool ordering, retrieved context, and previous state.

AI doesn't necessarily make systems simpler. Good architecture decides where complexity is actually worth paying for.

Collapse
 
hannune profile image
Tae Kim •

The SQL vs vector search one got us embarrassingly late in a project. We'd already wired up embeddings and were chasing why recall felt inconsistent, and it took someone from outside the team pointing out that the query was literally just "orders for this customer in this date range." Switched it to a plain query and the problem went away immediately. In hindsight it was obvious but we were deep in the AI pipeline mindset and couldn't see it.

Collapse
 
michaelhairetis profile image
Michael Hairetis •

my if statements were converted to ai if statements - a neat useful trick when you can manage it

Collapse
 
cyclopt_dimitrisk profile image
Dimitris Kyrkos •

That is a fair point because sometimes the condition itself is fuzzy and cannot be easily hardcoded. Using an LLM as a soft router for unstructured inputs is a solid pattern, provided you keep that step isolated so the rest of your application logic remains predictable.

Some comments have been hidden by the post's author - find out more