DEV Community

Which AWS limit is actually current? An agent that proves it, 32 vs 5 vs 16

Sarvar Nadaf on October 01, 2026

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content What I Built Which AWS quota is actual...
Collapse
 
salman_khan_c31307505285e profile image
Salmankhan •

Much better way to use Ai and limitations were sorted out in better way.

Picked as gem
Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

Exactly 💯

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

The rule the model is not allowed to override is the whole design — everything else is decoration. An agent that can argue with its own sanity layer will eventually win that argument. Showing the replaced value next to the live one is the other half: a number without its predecessor is unfalsifiable. One edge worth poking: quota changes propagate unevenly across regions, so the live API can lag the console for a while. Does the agent timestamp its proof and flag API-vs-console disagreement instead of picking a winner?

Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

"an agent that can argue with its own sanity layer will eventually win that argument" is a better sentence than anything in my post, im stealing that framing. and youre right about the predecessor, a number without the one it replaced is unfalsifiable, thats the half people drop.

on your edge: honest answer, no. the agent does not currently timestamp the live proof or flag api-vs-console as a disagreement, it folds both into one DRIFT verdict, which is too blunt. you and arhan both landed on the same soft spot from different angles. regional propagation lag is a real case i didnt handle, the live call is a point-in-time read with no "as of" stamp on it. the right fix is to stop calling it DRIFT and instead report "console says X, live api says Y, as of " and let that stand as the finding instead of resolving it. that'd be the first thing i change.

Collapse
 
sinarezaei profile image
Sina Rezaei •

This is a really solid approach to a problem that looks simple until you actually have to trust the answer.

I especially like that the model is not allowed to decide which number is correct. The flow is much more interesting: structured facts → source precedence → deterministic reconciliation → model response → guard → live AWS check.

The 32 vs 5 vs 16 example makes the point really well. Instead of hiding the conflict and returning one confident number, the system exposes the disagreement and then checks whether the reconciled value has drifted from the live environment.

That pattern is useful far beyond AWS. I can see the same architecture working for API versions, pricing, compliance rules, dependency versions, or even internal company policies where the question is basically: “Which value is actually current?”

For me, the strongest part isn't the AI agent itself. It's the architecture around the agent that limits what the AI is allowed to claim. That's a much more practical way to build reliable AI systems.

Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

thanks, you read it exactly how i hoped someone would. that line, "its not the agent, its the architecture that limits what the agent is allowed to claim" is the actual thesis and most people skim past it to look at the ai part. the guard is the whole project. the agent could be gpt or nova or a markov chain, doesnt matter, it still cant invent the number. and yeah the pattern is domain-agnostic on purpose, api versions and dependency versions were the two i kept thinking "this is the real use case" about. aws was just the demo where i could get three official pages to disagree on camera.

Collapse
 
arhancanli profile image
Arhan Canli •

The TF-IDF control that returns 16,000 is a great way to meet the "would keyword search get it?" bar: it shows the failure instead of asserting it.

One thing worth separating in the EC2 row, because it changes what DRIFT means. The live Service Quotas value for an account is the applied quota, which includes any increase that account was granted, so 16 vs 5 can be "this account asked for more" rather than "the record is stale". Service Quotas exposes both sides: GetServiceQuota gives the applied value and GetAWSDefaultServiceQuota gives the default. Comparing the record to the default tells you whether the docs are out of date; comparing the default to the applied value tells you about this account. Mixing them into one DRIFT flag will raise false alarms on any account that has had an increase.

The RDS row has a similar twist: describe-db-engine-versions on a live account can still list majors that are past standard support but in extended support, so a live 11 doesn't necessarily contradict a record saying 13 is the oldest standard-support major. A field for which support tier each number refers to would make that row read as AGREE-on-different-questions rather than drift.

Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

this is the best comment on the post and youre right, i'm conflating two different questions into one DRIFT flag.

to be precise about what the agent actually does today: verify_live calls GetServiceQuota, so its reading the applied value, which as you say already includes any increase this account was granted. so the EC2 "5 record vs 16 live" row is not clean evidence the record is stale, it could just as easily be "this account got an increase to 16". i presented it as drift and it isnt necessarily. thats a real bug in the framing, not a nuance.

the fix is exactly what you laid out: pull both, compare record-to-default (GetAWSDefaultServiceQuota) to answer "are the docs out of date", and default-to-applied (GetServiceQuota) to answer "what did this account change". two separate signals, never one flag. that also kills the false alarm on any account thats ever had an increase, which is most real accounts.

and the rds row is the same disease in a different organ, i checked after your comment: describe-db-engine-versions on a live account will happily list 11 because its in extended support, not because 13 isnt the oldest standard-support major. a live 11 and a record of 13 are answering two different questions ( whats runnable vs whats in standard support) and calling that drift is wrong. a support-tier field on the fact would make it read as "agree, different questions" like you said.

i'm not going to pretend i designed around this, i didnt, the demo account just happened to make the numbers look like a tidy drift story. the deterministic reconcile between the recorded sources is solid, the live layer needs to split applied/default/support-tier before any of its verdicts are trustworthy. genuinely thank you, this is the kind of correction that makes the thing real instead of a demo. did you hit this the hard way on a real account, a false alarm after a quota increase?

Collapse
 
arhancanli profile image
Arhan Canli •

Glad it helped, and credit to you for publishing the record and the live check side by side; that's what made the gap visible. To answer your question: no, not the hard way. It came from how Service Quotas reports the two values, not from a false alarm of my own. One more field worth keeping with each live value is the region: Service Quotas answers per region, so an increase in one region and the default in another can both be "current" for the same account. With region, applied vs default, and support tier on each fact, every row says exactly which question it answers.

Collapse
 
rulestack profile image
Rulestack •

Ours was Claude Code's limit on MCP tool output. The docs page gives 25,000 tokens. When we measured it in September on v2.1.273, the token count only ran once a result was already long in characters (somewhere between 45,000 and 52,000), so 24,000 characters of CJK text went in uncounted and grew the next request by 49,964 tokens, about twice the documented figure. Unlike your 16,000 row, the page wasn't out of date; it just doesn't say the limit is only checked once a result passes a certain length in characters.

Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

oh that is a nastier bug than mine, because mine was just stale and yours is the documented limit being technically true but measured in the wrong unit at the wrong time. 24k chars of CJK going in uncounted and then detonating the next request by ~50k tokens is brutal, and the worst part is the page isnt wrong so you'd never think to doubt it. thats actually a cleaner example of my whole point than my 16,000 row is: the failure isnt "the docs lied", its "the docs didnt tell you the condition under which the number applies". a char-length gate before a token count is exactly the kind of hidden precondition structured content is supposed to surface. did you end up measuring the real threshold or just clamping output well under it to be safe?

Collapse
 
micheypico profile image
Micheal Heypico •

This matches what we see operating a model-routing layer (32 models, one key at heypico.ai): the deterministic scaffolding around the LLM is what makes multi-model setups viable. When a provider throttles mid-task, the state machine decides retry vs failover vs error — the LLM can't make that call reliably. Debugging a 'flaky agent' is usually debugging a missing state machine around a fine model.

Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

yeah this is exactly it. the model was never the hard part, the state machine around it was. in our case the "flaky agent" bug was nova dropping a tool argument mid-run, and the fix wasnt a better prompt, it was caching the last real fetch server-side so the model physically cant relay a bad value. the llm writes the sentence, the scaffolding owns the number. retry-vs-failover is the same shape of problem, just with providers instead of facts. 32 models on one key sounds like a lot of state to babysit. whats been the nastiest failure mode you hit, provider throttling or silent output drift?

Collapse
 
henry786 profile image
Henry •

Handling data drift between official docs and live systems is such a headache! At The Printing World, keeping material specs synced across inventory, pricing, and suppliers is tricky. Seeing structured content handle these discrepancies is super clever. Great project!