DEV Community

Sarvar Nadaf
Sarvar Nadaf Subscriber

Posted on AI-assisted

Which AWS limit is actually current? An agent that proves it, 32 vs 5 vs 16

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

Which AWS quota is actually current? You copied a default vCPU limit off the AWS docs, shipped it, and it broke in production. I have done this. The number was stale, and nothing warned me.

Here is what makes it nasty. For one AWS fact, three official pages can each give you a different number. An old User Guide says one thing. The Service Quotas console says another. A pricing page says a third. All look official. None of them tells you which is live today.

So I built an agent that answers "which AWS value is current?" and then proves it. It reads typed facts from a Sanity Context Knowledge Base, picks the winner with a deterministic rule the model is not allowed to override, and shows both the current value and the one it replaced, each with its source. Then it does the part most content agents skip: it calls the live AWS API (read-only) to check whether even the reconciled record is still true. On the EC2 vCPU quota it turns up three different numbers, docs 32, record 5, live 16, and puts all three on the screen.

Stack: Amazon Nova Pro on Bedrock, Strands Agents (Python), a Sanity Context Knowledge Base read over the hosted Context MCP, and a read-only live AWS cross-check.

Why a keyword search returns the wrong AWS value

The challenge sets a hard bar: "the strongest submissions show an agent that only works because the content was structured. If a keyword search would have gotten you the same answer, aim higher."

Fair. So I built a fact where keyword search gets it wrong on purpose. Real one too: the max IOPS per volume for an EBS general purpose SSD. The current number is 80,000 (gp3, from the EBS console). The old number is 16,000, the legacy gp2 ceiling, still sitting in an older SSD guide. That old guide repeats the exact phrase you would type into a search box: "maximum IOPS per volume for a general purpose SSD EBS volume." It is stale and wordy, so keyword scoring loves it.

Watch it pick the wrong answer. Same content, real TF-IDF, the question a person would actually ask:

Keyword/TF-IDF baseline for: "maximum IOPS per volume general purpose SSD"

  0.8626  EBS/limit: Maximum IOPS per volume for a general purpose SSD ... = 16000   <-- WRONG (old)
  0.2026  EBS/limit: Maximum provisioned IOPS per gp3 volume            = 80000   <-- right, ranked lower
  0.0218  S3/price:  S3 Standard storage, first 50 TB / month           = 0.023
  0.0185  EC2/quota: Running On-Demand Standard ...                      = 5
  ...
Enter fullscreen mode Exit fullscreen mode

The stale 16,000 wins by a mile. Keyword search hands you the wrong number and sounds sure of it.

Keyword search ranks 16000 first; the agent reconciles to 80000

Demo

Run the credential-free core yourself in under a minute (public dataset over anonymous GROQ, local reconcile, read-only live AWS check, no model or token needed):

# 1. clone + install
git clone https://github.com/simplynadaf/aws-source-of-truth-agent
cd aws-source-of-truth-agent
pip install -r requirements.txt

# 2. the credential-free path (public dataset + local reconcile + live AWS check)
python -m agent.reconcile_offline --service EBS --type limit --region us-east-1
# -> Verdict: 80000 IOPS (serviceQuotasConsole wins over officialDocs)

# 3. the keyword control that returns the WRONG answer:
python -m agent.baseline "maximum IOPS per volume general purpose SSD"
Enter fullscreen mode Exit fullscreen mode

For the full LLM agent through the Context MCP, copy .env.example to .env and add a Bedrock region plus a Sanity Context Viewer token (the README walks through it). There is a JSON mode too, for piping into something else:

python -m agent.ask --json "What is the maximum IOPS per volume for a gp3 EBS volume?"
Enter fullscreen mode Exit fullscreen mode

The real run, Nova Pro, us-east-1, read-only throughout:

Fact Reconciled (from the record) Superseded Live AWS Result
EC2 On-Demand Standard vCPU quota 5 (console) 32 (old user guide) 16 DRIFT
EBS gp3 max IOPS per volume 80,000 (console) 16,000 (old SSD guide) unavailable trusted (no such live quota)
S3 Standard $/GB-mo 0.023 (pricing page) 0.021 (stale blog) 0.023 AGREE
RDS PostgreSQL oldest major 13 (release notes) 11 (old tutorial) 11 DRIFT
Lambda concurrent executions 1000 (dev guide) n/a unavailable trusted
Graviton4 (R8g) availability Available (instance types) n/a Available AGREE

Look at the EC2 row. Docs say 32. The reconciled record says 5. The live account says 16. Three numbers for one quota, and the agent shows all three with their sources instead of picking one and hoping. The drift is not a failure. It is the honest answer: this is the current record, and here is where reality has already moved past it.

The agent's cited EC2 answer with the live-drift line and tool trail

Code

🌊 AWS Source of Truth

When your AWS docs, pricing page, and Service Quotas console disagree, an agent that knows which one is telling the truth, then checks the live API to see if even that record has drifted.

Sanity Challenge Sanity Context AWS Nova Pro Strands

Live Demo Read the Article

Stars Forks Issues


Sanity project id: 0q5ohtvv · ⭐ If a stale AWS number has ever bitten you in production, give this a star.

The Problem • Why Search Fails • How Structure Fixes It • The Twist • Getting Started • FAQ

📖 Table of Contents

🤔 The Problem

You copied a limit straight out of the AWS docs…

Full source, plus the credential-free path judges can run with no token.

How I Used Sanity

What I pointed Sanity Context at: my own Sanity content. Every fact is a typed awsFact document in the production dataset, not a wall of prose:

{
  "_type": "awsFact",
  "service": "EBS", "factType": "limit", "region": "us-east-1",
  "key": "Maximum provisioned IOPS per gp3 volume",
  "currentValue": "80000", "unit": "IOPS",
  "effectiveDate": "2026-01-15",
  "source": { "name": "EBS gp3 volume limits (Service Quotas / EBS console)",
              "kind": "serviceQuotasConsole", "url": "..." }
}
// ...plus a separate record for the old 16,000 (kind: officialDocs, 2020).
Enter fullscreen mode Exit fullscreen mode

source.kind and effectiveDate are real fields, so the rule that picks the winner is dull and readable:

highest source precedence wins (console and pricing page beat changelog, which beats official docs, which beats a blog), and the newest effectiveDate breaks ties.

Prose cannot do this. The schema is the whole trick.

The Knowledge Base: a Sanity Context Knowledge Base indexes these facts. The build reads them ahead of time and writes short cited entries. Where two records fight, the entry keeps both numbers and both sources next to each other.

Which Context tools the agent used: the agent pulls facts over the hosted Context MCP with knowledge_base_search (ranked lookup) then knowledge_base_read (full entry). That is the proof the answer came from Sanity and not from the model's memory. (It can also fall back to anonymous GROQ over the public dataset for the no-token path.)

The Knowledge Base in the Sanity Context dashboard

The ebs/quotas_and_limits entry: current 80,000 vs superseded 16,000, with sources

What the agent does with what it retrieves: Nova Pro runs the tools and writes the sentence, but it does not get to decide the number. A plain function reconciles the value, and a guard checks the model's answer against it. If they disagree, the guard throws out the model's prose and ships the deterministic answer instead. On one EBS run Nova muddled its own wording, the guard caught it, and swapped in the correct answer with no help from me. The model cannot invent the number even if it tries.

The tool trail proves the path through Sanity on every question:

- [Sanity Context MCP] knowledge_base_search('EC2 quota us-east-1 ...') -> top='ec2/quotas_and_limits'; knowledge_base_read(['ec2/quotas_and_limits'])
- fetch_candidate_facts(service='EC2', fact_type='quota', region='us-east-1') -> 2 rows
- reconcile_facts(n=2) -> current=5
- verify_live(EC2/quota) -> drift
Enter fullscreen mode Exit fullscreen mode

Then: what if even the reconciled record is stale? (live AWS drift)

Reconciling the sources gives you the best answer the documents can offer. But documents rot. So the agent does one more thing a pure content agent will not: it calls the live AWS API, read-only, and asks whether the reconciled value is still true right now. Service Quotas, the Price List API, EC2, RDS. The results are in the Demo table above. The live layer is additive and clearly labelled; it is not part of Sanity and it is account-specific.

What didn't work (the honest part)

  • I planned to screenshot a resolved conflict in the dashboard. The build never gave me one. The Knowledge Base read the typed records, reconciled the 32-vs-5 fight into a clean entry with both numbers cited, and left the Issues queue empty. That is the KB doing its job, but it means there is no "I resolved an Issue" screenshot to show, so I am not claiming one.
  • Nova dropped a tool argument on me. fetch_candidate_facts returned a row, then Nova passed an empty string to reconcile_facts, which saw zero facts. The guard fail-closed to "Not verified" instead of guessing. I fixed it by caching the last real fetch server-side, so the facts always come from Sanity, never from whatever the model relayed.
  • A wrong factType guess used to lose a real fact. Nova guessed instanceType when the field value was regionalAvailability, the typed filter matched nothing, and the fact vanished. Now a zero-row typed fetch retries on service and region alone.
  • The live check is not part of Sanity, and it is account-specific. A judge on a different AWS account will see different live numbers. That layer is additive and labelled as such. Two facts (EBS max IOPS, Lambda) have no matching Service Quotas entry, so the agent says "unavailable" rather than fake a check. The seed is labelled demo data at the top.

Reusing the pattern beyond AWS

Drop the AWS parts and the shape is generic: type your sources, index them into a Knowledge Base, reconcile by precedence and date, guard the answer so the model cannot fake the winner, and check a live system of record when one exists. It fits API version support, pricing, compliance clauses, legal terms. Anywhere the real question is "which version is current?"

The thing I took away: the win here was not a smarter model. It was structured content, plus the nerve to say "the record says 5, but reality says 16."

Sanity Project Details

If your docs and your console have ever disagreed with each other, tell me which number bit you.

Top comments (14)

Collapse
 
salman_khan_c31307505285e profile image
Salmankhan •

Much better way to use Ai and limitations were sorted out in better way.

Picked as gem
Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

Exactly 💯

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

The rule the model is not allowed to override is the whole design — everything else is decoration. An agent that can argue with its own sanity layer will eventually win that argument. Showing the replaced value next to the live one is the other half: a number without its predecessor is unfalsifiable. One edge worth poking: quota changes propagate unevenly across regions, so the live API can lag the console for a while. Does the agent timestamp its proof and flag API-vs-console disagreement instead of picking a winner?

Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

"an agent that can argue with its own sanity layer will eventually win that argument" is a better sentence than anything in my post, im stealing that framing. and youre right about the predecessor, a number without the one it replaced is unfalsifiable, thats the half people drop.

on your edge: honest answer, no. the agent does not currently timestamp the live proof or flag api-vs-console as a disagreement, it folds both into one DRIFT verdict, which is too blunt. you and arhan both landed on the same soft spot from different angles. regional propagation lag is a real case i didnt handle, the live call is a point-in-time read with no "as of" stamp on it. the right fix is to stop calling it DRIFT and instead report "console says X, live api says Y, as of " and let that stand as the finding instead of resolving it. that'd be the first thing i change.

Collapse
 
sinarezaei profile image
Sina Rezaei •

This is a really solid approach to a problem that looks simple until you actually have to trust the answer.

I especially like that the model is not allowed to decide which number is correct. The flow is much more interesting: structured facts → source precedence → deterministic reconciliation → model response → guard → live AWS check.

The 32 vs 5 vs 16 example makes the point really well. Instead of hiding the conflict and returning one confident number, the system exposes the disagreement and then checks whether the reconciled value has drifted from the live environment.

That pattern is useful far beyond AWS. I can see the same architecture working for API versions, pricing, compliance rules, dependency versions, or even internal company policies where the question is basically: “Which value is actually current?”

For me, the strongest part isn't the AI agent itself. It's the architecture around the agent that limits what the AI is allowed to claim. That's a much more practical way to build reliable AI systems.

Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

thanks, you read it exactly how i hoped someone would. that line, "its not the agent, its the architecture that limits what the agent is allowed to claim" is the actual thesis and most people skim past it to look at the ai part. the guard is the whole project. the agent could be gpt or nova or a markov chain, doesnt matter, it still cant invent the number. and yeah the pattern is domain-agnostic on purpose, api versions and dependency versions were the two i kept thinking "this is the real use case" about. aws was just the demo where i could get three official pages to disagree on camera.

Collapse
 
arhancanli profile image
Arhan Canli •

The TF-IDF control that returns 16,000 is a great way to meet the "would keyword search get it?" bar: it shows the failure instead of asserting it.

One thing worth separating in the EC2 row, because it changes what DRIFT means. The live Service Quotas value for an account is the applied quota, which includes any increase that account was granted, so 16 vs 5 can be "this account asked for more" rather than "the record is stale". Service Quotas exposes both sides: GetServiceQuota gives the applied value and GetAWSDefaultServiceQuota gives the default. Comparing the record to the default tells you whether the docs are out of date; comparing the default to the applied value tells you about this account. Mixing them into one DRIFT flag will raise false alarms on any account that has had an increase.

The RDS row has a similar twist: describe-db-engine-versions on a live account can still list majors that are past standard support but in extended support, so a live 11 doesn't necessarily contradict a record saying 13 is the oldest standard-support major. A field for which support tier each number refers to would make that row read as AGREE-on-different-questions rather than drift.

Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

this is the best comment on the post and youre right, i'm conflating two different questions into one DRIFT flag.

to be precise about what the agent actually does today: verify_live calls GetServiceQuota, so its reading the applied value, which as you say already includes any increase this account was granted. so the EC2 "5 record vs 16 live" row is not clean evidence the record is stale, it could just as easily be "this account got an increase to 16". i presented it as drift and it isnt necessarily. thats a real bug in the framing, not a nuance.

the fix is exactly what you laid out: pull both, compare record-to-default (GetAWSDefaultServiceQuota) to answer "are the docs out of date", and default-to-applied (GetServiceQuota) to answer "what did this account change". two separate signals, never one flag. that also kills the false alarm on any account thats ever had an increase, which is most real accounts.

and the rds row is the same disease in a different organ, i checked after your comment: describe-db-engine-versions on a live account will happily list 11 because its in extended support, not because 13 isnt the oldest standard-support major. a live 11 and a record of 13 are answering two different questions ( whats runnable vs whats in standard support) and calling that drift is wrong. a support-tier field on the fact would make it read as "agree, different questions" like you said.

i'm not going to pretend i designed around this, i didnt, the demo account just happened to make the numbers look like a tidy drift story. the deterministic reconcile between the recorded sources is solid, the live layer needs to split applied/default/support-tier before any of its verdicts are trustworthy. genuinely thank you, this is the kind of correction that makes the thing real instead of a demo. did you hit this the hard way on a real account, a false alarm after a quota increase?

Collapse
 
arhancanli profile image
Arhan Canli •

Glad it helped, and credit to you for publishing the record and the live check side by side; that's what made the gap visible. To answer your question: no, not the hard way. It came from how Service Quotas reports the two values, not from a false alarm of my own. One more field worth keeping with each live value is the region: Service Quotas answers per region, so an increase in one region and the default in another can both be "current" for the same account. With region, applied vs default, and support tier on each fact, every row says exactly which question it answers.

Collapse
 
rulestack profile image
Rulestack •

Ours was Claude Code's limit on MCP tool output. The docs page gives 25,000 tokens. When we measured it in September on v2.1.273, the token count only ran once a result was already long in characters (somewhere between 45,000 and 52,000), so 24,000 characters of CJK text went in uncounted and grew the next request by 49,964 tokens, about twice the documented figure. Unlike your 16,000 row, the page wasn't out of date; it just doesn't say the limit is only checked once a result passes a certain length in characters.

Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

oh that is a nastier bug than mine, because mine was just stale and yours is the documented limit being technically true but measured in the wrong unit at the wrong time. 24k chars of CJK going in uncounted and then detonating the next request by ~50k tokens is brutal, and the worst part is the page isnt wrong so you'd never think to doubt it. thats actually a cleaner example of my whole point than my 16,000 row is: the failure isnt "the docs lied", its "the docs didnt tell you the condition under which the number applies". a char-length gate before a token count is exactly the kind of hidden precondition structured content is supposed to surface. did you end up measuring the real threshold or just clamping output well under it to be safe?

Collapse
 
micheypico profile image
Micheal Heypico •

This matches what we see operating a model-routing layer (32 models, one key at heypico.ai): the deterministic scaffolding around the LLM is what makes multi-model setups viable. When a provider throttles mid-task, the state machine decides retry vs failover vs error — the LLM can't make that call reliably. Debugging a 'flaky agent' is usually debugging a missing state machine around a fine model.

Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

yeah this is exactly it. the model was never the hard part, the state machine around it was. in our case the "flaky agent" bug was nova dropping a tool argument mid-run, and the fix wasnt a better prompt, it was caching the last real fetch server-side so the model physically cant relay a bad value. the llm writes the sentence, the scaffolding owns the number. retry-vs-failover is the same shape of problem, just with providers instead of facts. 32 models on one key sounds like a lot of state to babysit. whats been the nastiest failure mode you hit, provider throttling or silent output drift?

Collapse
 
henry786 profile image
Henry •

Handling data drift between official docs and live systems is such a headache! At The Printing World, keeping material specs synced across inventory, pricing, and suppliers is tricky. Seeing structured content handle these discrepancies is super clever. Great project!