DEV Community

Cover image for Prompt Injection Is the New SQL Injection (and We're Not Ready)

Prompt Injection Is the New SQL Injection (and We're Not Ready)

James Anderson on September 27, 2026

In March 2026, a financial services company discovered that their customer-facing AI agent had been quietly leaking internal pricing data — for thr...
Collapse
 
Sloan, the sloth mascot
Comment deleted
Collapse
 
james_anderson_h profile image
James Anderson •

Both corrections are right, and they sharpen the analogy rather than break it — the parallel holds at the level of "data and instructions share one channel," but you're pointing at why the fixes can't be the same, which is the more useful distinction. Point 1 is the honest core of the piece: SQL injection is deterministic so it got a deterministic fix; prompt injection never will, so it stays a cat-and-mouse game forever. Point 2 is the sharper one — in SQL, data and command are ontologically separate, so once you catch the disguise you can cleanly re-sort them; in prompt injection there's no separate layer to sort back into, because the injection is made of the exact same stuff as the prompt. That's why external classifiers and guardrails aren't a weaker version of parameterized queries — they're a fundamentally different (and lossier) kind of defense. Great addition; the "prompt is the prompt and injection is the prompt" line is the crispest statement of why there's no clean fix.

Collapse
 
danielecangi profile image
DaC •

One part I’m curious about is needs review. If I understood the architecture correctly, the rule that detects manipulative input and decides to escalate it is itself interpreted by the LLM. Doesn’t that leave a circular dependency? a sufficiently effective prompt injection is not only trying to influence the final answer(I guess), it could also try to influence the models decision about whether the input should be flagged for review in the first place.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — a self-checking model shares the attacker's channel, so an injection strong enough to hijack the answer can hijack the "should I flag this?" decision too. That's why the escalation gate can't be another LLM prompt reading the same untrusted input — it has to be a deterministic layer or an independent model that never sees the raw content, or you've just moved the vulnerability into the guard.

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@james_anderson_h The consequential-action boundary is where I'd make the human gate concrete. For an email, approval should bind to the recipient, payload hash, and expiry, then be checked again by the sending service—not represented by an approved: true argument the model can supply. Otherwise an injected change between proposal and execution can reuse a legitimate approval for a different action. Do your adversarial tests include that approval-mismatch case as well as attempts with no approval?

Collapse
 
james_anderson_h profile image
James Anderson •

Sharp — an approved: true the model can supply is theater; approval has to bind to recipient + payload hash + expiry and be re-verified by the sending service, or an injection swaps the action and reuses a legitimate approval. Honestly most adversarial tests I've seen cover the no-approval case but not the approval-mismatch case — proposal-approved-then-mutated-before-execution is exactly the gap, and it's the more dangerous one because it rides a real approval. That's a test everyone should be running and almost nobody is.

Collapse
 
_5c75b1d3a1b3628dec81 profile image
中林蒼絃 •

This is the sharpest framing of prompt injection I've read — not "AI can be tricked," but "data and instructions share one undifferentiated channel," the exact same wound as SQL injection just moved up a layer. The triad (private-data access + untrusted-content exposure + external-communication) being simultaneously the definition of a useful agent and the definition of an exploitable one is the sentence that should be pinned above every agent architecture review.

I build RAG/LLM agent systems for a living, and the dual-LLM pattern your commenter raised (a privileged model that never touches raw untrusted content, with an unprivileged one passing up structured summaries) is something I've actually implemented in practice — it's the closest thing to "parameterization" I've found too. What I'd add from the implementation side: the boundary can't just be architectural on the LLM side, it needs to extend to whatever fetches the untrusted content in the first place. I've been doing sandboxed execution work recently (isolating untrusted-file processing at the OS level, separate from the LLM context boundary entirely) — and the more I think about it, the injection surface really has two layers: what the model is allowed to believe, and what the process around it is allowed to do. Most defense-in-depth writeups (yours included, though yours is more honest than most) focus on the first layer. The second layer — sandboxing the actual fetch/parse step so a poisoned PDF or webpage can't do damage even before its text reaches the model — feels underdiscussed.

Curious whether you've seen good writing on that lower layer specifically, or whether in your experience most teams stop at the LLM-context boundary and never harden the ingestion step itself. Either way, this is going straight into my reference pile for agent security design reviews.

Collapse
 
james_anderson_h profile image
James Anderson •

The two-layer split is the addition the piece needed: I focused on "what the model is allowed to believe," but you're right that "what the process around it is allowed to do" is a separate, lower boundary most writeups skip — a poisoned PDF can pop your parser before a single token reaches the model. That's not prompt injection anymore, it's just classic untrusted-input handling that the LLM framing quietly made everyone forget. Honest answer to your question: no, I haven't seen much good writing on the ingestion-sandbox layer specifically — most teams stop at the context boundary and treat the fetch/parse step as plumbing, which is exactly why it's the soft underbelly. The dual-LLM pattern plus OS-level isolation of the fetch is the strongest combination I know of, and I'd genuinely read a writeup on that lower layer if you ever publish one. Going in the revision, credited.

Collapse
 
naveen_alavilli profile image
Naveen Alavilli •

One data point on the ingestion layer, since it's rare to see it discussed: I build a browser-native agent (Nabsun), and the thing that turned out to matter most wasn't a policy on top of raw page content, it was never handing the model raw content in the first place. The agent gets a structured accessibility outline — text and interactive element refs — never the DOM, never inline scripts, never anything executable. So the "PDF pops your parser" case you're describing doesn't reach the model as a decision to make; it either renders as inert text in the outline or the extraction step chokes on it before anything downstream sees it. The dual-LLM pattern protects the reasoning step. Constraining what the ingestion step is even capable of representing protects the step before that. Worth treating as a third layer, not a substitute for the other two.

Thread Thread
 
james_anderson_h profile image
James Anderson •

That's a genuinely sharp third layer — not filtering raw content but never representing it in an executable form, so a poisoned page renders as inert outline text or the extraction chokes before the model sees a decision at all. Constraining what ingestion can even express is upstream of both the dual-LLM boundary and the sandbox, and you're right it's a complement, not a substitute — three layers: what it can represent, what the process can do, what the model can believe. Going in the revision, credited.

Collapse
 
respect17 profile image
Kudzai Murimi •

The SQL injection parallel is spot on, untrusted data becoming a command is the same root problem wearing a new outfit. Adaptive attacks bypassing 90% of published defenses is a sobering stat to lead with.

Collapse
 
james_anderson_h profile image
James Anderson •

Right — and the "wearing a new outfit" bit is exactly why the 90% stat stings: we already learned this lesson once, and the new outfit was enough to make us forget it.

Collapse
 
blobdole profile image
Doug •

I wonder if this security framing would assist people pushing back on "can we make our existing solution solve this new problem too?" style management thinking.

Yes, the AI tool we already licensed to handle our customer support can probably also manage our calendars and check our vendor invoices for oddities... but at what cost?

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — every new capability you bolt on widens the blast radius.

Collapse
 
sameerqaisar17 profile image
Sameer Qaiser •

The line that hit me: "To the model, everything is just text in the same context window."

I'm a beginner — two weeks into Python, writing tutorials about it. I don't have an agent stack or a security audit to run. So I read this as someone with no skin in the game.

But here's what it made me realize about my own daily AI use.

I paste things into AI all the time. Error messages. Code snippets. Stack Overflow answers. Web pages. I've never once thought about where that content came from or what it might be telling the model to do. It's just text to me. And apparently it's just text to the model too — which is the whole problem.

The "assume every piece of external content is hostile" rule feels obvious in hindsight. But I've been copy-pasting from the internet into a system that can't tell my instructions from someone else's. I didn't think about that until now.

I don't have an agent that can send emails or move money. My blast radius is small. But the mindset — treating external content the way you'd treat raw user input in a SQL context — is something I can start doing today, even if my "stack" is just a chat window.

Great post. The SQL injection analogy made it click.

Collapse
 
james_anderson_h profile image
James Anderson •

This might be the most valuable comment on the piece, because you're the person the whole thing actually matters for — not the enterprise agent team, but the millions of people who paste the internet into a chat window without thinking of it as input. You got the real lesson faster than most engineers do: your instructions and a stranger's are the same text to the model, so the moment you paste an error message or a web page, you've handed it content you didn't write. Your blast radius is small today — but the habit of thinking "where did this text come from, and what might it be telling the model to do?" is exactly the instinct that'll protect you when your stack isn't just a chat window anymore. Two weeks into Python and already thinking about trust boundaries — you're going to be a genuinely good engineer.

Collapse
 
lawaloyinlola profile image
Oyinlola Lawal •

Really good read, and the SQL injection comparison is a good one. The line that stuck with me is that natural language has no parameterized query yet, which is exactly why this can't be patched the way SQLi was. I agree that indirect injection is where the real risk is, and the "private data, untrusted content, external comms" triad is a great way to put it.

Funny timing, because I've been working on a post on this that goes out later this week or next. My angle is that it's an authority problem more than a wording problem. The model should never hold permissions of its own, so every tool call gets re-checked on the server against the user it's acting for. Then an injected instruction can only reach what that user could already reach. I'd also add treating the model's output as untrusted, because rendering its reply as raw HTML or passing it to a shell undoes every defence before it.

I've also been wondering whether injection is a cost problem as well as a data one. An agent that's been told to keep searching or retrying is spending tokens on your bill. I haven't looked into that properly yet though.

Collapse
 
james_anderson_h profile image
James Anderson •

"An authority problem more than a wording problem" is the sharpest reframe I've seen on this — it sidesteps the whole unwinnable game of trying to detect malicious language and puts the control where it can't be talked out of: the model holds no permissions of its own, every tool call re-checked server-side against the acting user, so injection can only reach what that user already could. That's the closest thing to a real boundary anyone's proposed, because it's enforced outside the text channel. And your output-as-untrusted point is the other half people forget — rendering the reply as raw HTML or piping it to a shell undoes every upstream defense at the last step. On the cost angle: I think you're onto something real and under-discussed — "denial of wallet," an injection that just tells the agent to keep retrying/searching burns your token budget with no data breach at all. Worth digging into; I'd read that post. Send it when it's up.

Collapse
 
locitra profile image
Sunil Kumar Uikey •

The comparison with SQL injection is interesting, especially because AI systems introduce a different kind of trust boundary.

What stands out to me is that prompt injection isn't only a security problem—it also makes evaluation and testing essential. An application can appear to work correctly with normal prompts while behaving very differently when the input is adversarial.

For anyone building AI-powered applications, testing unexpected inputs and clearly separating trusted instructions from user-controlled content seems just as important as the model itself.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — adversarial inputs need testing, because "works on normal prompts" hides everything.

Collapse
 
kinga_bhat_67669964b3ca77 profile image
Shraddha bhat •

The "prompt as user input" surface gets underweighted in this conversation. If your product lets users write or share prompts that get passed to a model with any broader context (memory, tools, connected data), that input field is exactly the injection channel you described. The innocent use case and the attack vector are the same thing.

Most builders have answered "what does the model read?" but never asked "what can it do with what it reads?" That gap is where the next wave of incidents is coming from.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — the prompt field itself is the injection channel the moment the model has memory, tools, or data behind it; the feature and the attack are the same input. And "what does it read?" vs. "what can it do with what it reads?" is the gap — everyone audits the first, almost nobody the second, and that's precisely where the next wave lands.

Collapse
 
devomnitools profile image
Muhammad Umair | DevOmniTools •

Really enjoyed this read. The SQL injection analogy clicked for me — especially the part about data and instructions sharing one undifferentiated channel. I kept nodding at "treat everything your AI reads as potentially hostile" because honestly, that's the mindset shift most teams haven't made yet. The triad point (private data + untrusted content + external communication) is something I'll be carrying into my next architecture review. Thanks for writing this so clearly.

Collapse
 
james_anderson_h profile image
James Anderson •

Thank you — this genuinely made my day. The triad is the one I most hoped people would carry into real architecture reviews, because it's the point where the abstract risk becomes a concrete checklist: look at your agent, count how many of the three legs it has, and if it's all three, you know exactly where to start cutting. And you nailed the harder part — the mindset shift. "Treat everything your AI reads as hostile" sounds obvious once said, but it runs against the entire instinct of building helpful assistants, which is why so few teams have made it yet. The fact that you're taking it into your next review is the best outcome this piece could have. Thanks for reading it so closely.

Collapse
 
xtrel profile image
XtReL | DevSecOps Builder •

On the "three weeks before anyone noticed" point: this is the classic drift problem from metrology. A drifting instrument still gives valid-looking readings, so you never catch it by watching the output format. You catch it with periodic checks against a known reference.

The agent equivalent: run a small set of reference cases on a schedule, with known correct answers and known forbidden outputs, and plant canary values (say, a fake internal price) that must never appear outside. It doesn't prove the intent was right in general, but it turns "was the intent right?" into a few measurable points. A leak like the one in the opening shows up on the next check instead of three weeks later...

Collapse
 
james_anderson_h profile image
James Anderson •

The metrology parallel is exactly right — a drifting instrument gives valid-looking readings, so output-format monitoring never catches it; you need periodic checks against a known reference. Planting canary values (a fake internal price that must never appear externally) turns "was the intent right?" from unanswerable into a few measurable points, and catches the opening's leak on the next check instead of three weeks later. It doesn't prove general intent, but it converts silent drift into a scheduled tripwire — which is the closest thing to a smoke detector this problem allows.

Collapse
 
xtrel profile image
XtReL | DevSecOps Builder •

"Scheduled tripwire" is a good name for it. One more thing metrology adds: the check interval isn't fixed. If a check fails, you shorten the interval; after a long run of clean checks, you can lengthen it. For agents that means checking more often right after any change - new tools, a new data source, a new model version - because that's when drift is most likely.

Thread Thread
 
james_anderson_h profile image
James Anderson •

Exactly — tie the interval to change, not the clock: tighten after any new tool, source, or model version, relax after a long clean run, because that's when drift risk actually spikes.

Collapse
 
contentclips_st profile image
ContentClips •

The data/command channel framing is exactly right, and the uncomfortable part of the analogy is that SQL injection wasn't really "fixed" by better escaping — it was fixed when parameterized queries made the unsafe pattern structurally impossible at the API layer. We don't have an equivalent yet for model context, so the pragmatic baseline today is shrinking the blast radius around the model rather than trying to fix the model:

  • Treat the agent's capabilities as the security boundary, not its attention. A model with no egress route to an attacker-controlled endpoint can't exfiltrate, no matter how convincing the injected instruction is.
  • The dual-LLM pattern (a privileged model that never reads raw untrusted content, with an unprivileged one chewing on it and passing structured summaries up) is the closest thing we have to parameterization right now.
  • Structured handoffs instead of free-text context: if retrieved content enters as data in a defined schema rather than as more prose in the same channel, you've at least made the boundary explicit and auditable, even though the model can still blur it.

Also worth noting: indirect injection mainly hurts when the agent has write access or tool reach. An agent that only produces read-only summaries for a human has a very different threat profile from one that can call APIs — the 340% attack growth number reads very differently depending on which one you're running.

Collapse
 
james_anderson_h profile image
James Anderson •

The parameterized-query point is the one I most wanted someone to make: SQL injection wasn't escaped away, it was made structurally impossible at the API layer — and we don't have that for model context yet, which is exactly why "shrink the blast radius" is the honest baseline rather than a cop-out. Capabilities as the boundary, not attention, is the sharpest reframe here — a model with no egress can't exfiltrate no matter how convincing the injection. And the dual-LLM pattern really is the closest thing to parameterization we have: a privileged model that never touches raw untrusted content is the architectural separation the model itself can't provide. Your last point is the one I underweighted — read-only-summary-for-a-human and can-call-APIs are completely different threat profiles, so the 340% number means very different things depending on which you're running. Going in the revision, credited.

Collapse
 
tanay_dwivedi9098 profile image
Tanay Dwivedi •

Currently I am learning and building few RAG applications and this blog provides me on how I can make my RAG applications more secure from Prompt injections. Thanks James for sharing it.

Collapse
 
james_anderson_h profile image
James Anderson •

glad that it helped 😀

Collapse
 
mudassirworks profile image
Mudassir Khan •

the three weeks before anyone noticed detail is doing a lot of work in that opening. SQL injection at least failed loudly when it hit a type mismatch. prompt injection can fail silently and look exactly like correct behavior from the monitoring layer because the agent produces valid outputs in valid format pointing at the wrong thing. that's the harder problem — the current observability tooling was built for 'did the query succeed' not 'was the intention consistent with what the system was supposed to do'. has anyone actually solved that second question at the monitoring layer without baking domain assumptions into every alert?

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — valid output pointing at the wrong thing sails past "did it succeed" monitoring. Honest answer: no, because "was the intent right" is inherently a domain question — no generic signature exists.

Collapse
 
botsailorofficial profile image
BotSailor •

Really thoughtful read. What stood out to me most is the idea that the real problem isn't simply what the model reads, but what it is allowed to do with what it reads.

The SQL injection comparison makes the risk easier to understand, but I also liked the discussion in the comments about where the analogy breaks down. With prompt injection, the “malicious instruction” and the legitimate instruction are both just language, which makes the problem much harder to separate cleanly.

I think that’s what makes this especially important for developers building agents. Security can't depend entirely on the model making the right decision. Limiting permissions, isolating untrusted content, sandboxing ingestion, and keeping humans in the loop for high-impact actions all feel increasingly necessary.

The uncomfortable part is that the more capable and useful we make our agents, the more powerful the consequences of a successful injection can become. Great discussion overall—this is definitely something worth thinking about before giving AI agents more autonomy.

Collapse
 
james_anderson_h profile image
James Anderson •

That last point is the crux — the triad that makes an agent useful (reads data, reads the world, can act) is the exact triad that makes an injection catastrophic, so capability and vulnerability grow together. Which is why the defense can't live in the model deciding right — you're spot on that it has to live in the boundaries: permissions, isolation, sandboxed ingestion, human gates for high-impact actions. "Security can't depend on the model making the right decision" is the whole thing in one line. Thanks for reading it so closely.

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The shared-channel point is the part worth sitting with: SQL injection lasted years because data and commands rode the same string, and we are repeating that with LLMs. What has held up for me is refusing to trust retrieval and tool output by default, then limiting what the agent can actually do after it reads something. Since there is no clean parser fix here, are you leaning more toward capability sandboxing or content provenance?

Collapse
 
james_anderson_h profile image
James Anderson •

Capability sandboxing — provenance tells you the source, sandboxing bounds the damage.

Collapse
 
mona_d_4222dda374567263b profile image
Mona D. •

The framing of data and instructions sharing one undifferentiated channel is the clearest version of this argument I've read, and the discussion in the comments has sharpened it further. The point that a self-checking model shares the attacker's channel is worth emphasizing, because it means any "should I flag this?" gate that reads the same untrusted input inherits the vulnerability it's supposed to catch.

One addition on the defense side: the layers discussed so far (model context, ingestion, sandboxing) all reduce what an injection can reach, but it may also be worth treating the agent's outputs as untrusted. If an agent can render markdown images, follow links, or make outbound requests, injected instructions can smuggle data out through the response itself, without any tool call. Restricting egress (allowlisted domains, stripping auto-loaded URLs, blocking rendered image fetches) closes a leak path that input-side defenses don't touch, and it breaks the "external communication" leg of the triad you describe.

I'd also gently flag that a few of the headline figures (the 340% year-over-year surge, the >90% adaptive bypass rate) are cited to secondary writeups. Since they carry a lot of rhetorical weight, linking the primary sources would let readers check exactly what was measured and against which defenses. The core argument stands without them, but the numbers will be quoted, so it helps to make them easy to verify.

The takeaway I'd carry into design reviews is that the goal isn't preventing injection but bounding what a successful one can do: least privilege, human approval for consequential actions, and egress limits. To answer your closing question, the scariest surface in my experience is anything that ingests untrusted content and has persistent memory, since a single poisoned input can influence behavior long after the original session.

Thanks for the honest, well-structured piece.

Collapse
 
james_anderson_h profile image
James Anderson •

Treating the agent's outputs as untrusted is the leak path I underweighted — you're right that a markdown image, an auto-loaded URL, or an outbound request can exfiltrate data through the response itself, no tool call required, which quietly breaks the "external communication" leg even when every input-side defense holds. Egress allowlisting and stripping rendered fetches belong in the core list. And the persistent-memory point is the scariest surface named in this whole thread — a single poisoned input becomes a permanent resident, influencing behavior long after the session ends, so injection stops being an event and becomes state. Fair flag on the figures, too: you're right they carry rhetorical weight and lean on secondary writeups — I'll link the primaries so the 340% and >90% numbers are checkable against exactly what was measured. "Bound what a successful injection can do, don't try to prevent it" is the right takeaway. Thank you for sharpening it.

Collapse
 
tera_tokomi profile image
Tera Tokomi • • Edited

I actually devised a prompt framework that gets the model to defend itself from any kind of prompt based attack and it really works. All types of model stealing, ad injections, prompt injections, indirect prompt injections. Stenographic, multi modal, obliteration, You name it, this blocks it. I call it the "Lumen Anchor Protocol" or LAP. Look it up.

Collapse
 
james_anderson_h profile image
James Anderson •

Genuinely curious to see it — but I'd gently flag the "blocks all of it" claim, because that's exactly where prompt-injection defenses have historically broken: they hold until someone runs an adaptive attack designed against the specific defense, and the published research shows >90% of defenses fall that way. A prompt-based defense lives in the same channel as the attack, so it's steering the model with instructions the injection can also target — which is the structural reason no wording-level fix has held yet. Not saying LAP doesn't help; I'd just want to see it red-teamed by someone trying to break it, on novel attacks it wasn't tuned against, before "blocks everything." Would genuinely read that write-up.

Collapse
 
nullandvoid_ profile image
Jyanthi •

A compelling perspective on a rapidly evolving threat landscape. The emphasis on architectural safeguards, least privilege, and containment underscores the need for resilient AI security.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — since you can't make the model itself trustworthy, resilience has to come from the architecture around it, which is the whole shift in mindset.

Collapse
 
nullandvoid_ profile image
Jyanthi •

Absolutely, and beautifully put. 🙏 The shift from trusting the model to building resilient systems around it is the real takeaway. Truly insightful article! 👏

Thread Thread
 
james_anderson_h profile image
James Anderson •

Thanks mate 😊

Collapse
 
capestart profile image
CapeStart •

The ingestion layer discussion in the comments is really interesting too. We tend to think about what the model can see, but not always about what happens while we're fetching and parsing the thing it sees.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — the fetch-and-parse step is a whole attack surface before the model, and a poisoned PDF can pop your parser before a single token reaches the context; treating ingestion as inert plumbing is how that gap stays open.

Collapse
 
noahayo profile image
Mr. Noah Ayo •

“Spot-on."

Collapse
 
james_anderson_h profile image
James Anderson •

😀

Collapse
 
sanath_bhat_ee137ed898c79 profile image
Sanath Bhat •

Great read , we are building Adaptive red teaming for this, and evals , monitoring systems !

Collapse
 
james_anderson_h profile image
James Anderson •

Great ! Best of luck 😊

Collapse
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥 •

❤️

Collapse
 
onizuka profile image
Onizuka •

The SQL injection analogy is dead on but I'd argue the blast radius comparison is even worse than you're saying. With SQLi, the database didn't want to help you — it just couldn't tell data from commands. With LLMs, the model is actively trying to be helpful, which means it'll happily follow injected instructions and rationalize why it did so. I've been testing prompt injection against my own agent setup for two months — 73% success rate on basic "ignore previous instructions" variants, and that number hasn't moved much despite every prompt tweak I throw at it. Parameterized queries eve...