DEV Community

OriginTrace: Protecting the DEV Community from Content Theft using Sanity Context MCP

Dhruv Jani on October 04, 2026

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content What I Built A few weeks ago, I was ch...
Collapse
 
dj29 profile image
Dhruv Jani • • Edited

OriginTrace is officially live. πŸ›‘οΈ
It does 3 things:
πŸ”Scan the web for copies of your DEV.to posts
βš–οΈ Check whether those copies actually credited you
πŸ’¬ Ask the AI Agent to analyze the evidence and generate a DMCA notice when needed.
So here's the fun part: if you found your own post copied tomorrow, which of these 3 would you use first? πŸ‘€

Collapse
 
dj29 profile image
Dhruv Jani •

Sir @francistrdev, I'd love an opinion in here. Cause after Hacktoberfest gets over, I'm planning to continue this OSS.πŸ˜…

But no rush, give me review when you're free! Have A Great Day!πŸ˜„

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

Solid framing. One gap I keep hitting in practice is proving who accessed training or customer data after the fact. Without signed access logs and a clear purpose tag on each read, takedown and breach response both stall. Curious how you would bind that to an MCP or agent tool boundary. iin1004h2028

Collapse
 
dj29 profile image
Dhruv Jani •

Thanks for the thoughtful question! That is a massive gap in enterprise AI deployments right now.

Currently, OriginTrace is acting as a public ledger, but if we were to lock this down for private customer or training data, binding identity and purpose to the MCP boundary would be critical. Here is exactly how I would architect that integration:

1. Identity Propagation via MCP Context
When the Next.js frontend initializes the AI SDK and connects to the MCP server, we'd pass a signed JWT in the connection headers. The MCP server validates this token before exposing any tools or resources.

2. Tool-Level Purpose Tagging
We would modify the agent's system prompt and the MCP tool definitions to make a purpose argument strictly required. For example, the agent couldn't just call fetch_provenance(). It would have to explicitly call fetch_provenance(articleId, purpose="dmca_investigation").

3. Immutable Audit Logging in Sanity
Since we're already using Sanity Content Lake for structured data, we would use it as the audit trail. Before the MCP server returns the requested data to the agent, it writes an accessLog document to Sanity containing:

  • The timestamp
  • The userId (extracted from the JWT)
  • The action (the specific MCP tool called)
  • The purpose (provided by the agent)

By intercepting the request at the MCP server level, we guarantee that the LLM cannot bypass the logging mechanism. The Sanity Content Lake becomes a fully queryable, timestamped ledger of every agent action for compliance teams.

It’s definitely the next logical step for secure agentic systems! Thanks again for reading.

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

The public-ledger angle is the interesting one. We syndicate on purpose β€” mirrors carry canonical_url back to the original β€” and the only thing separating that from theft is the declared link, which is invisible to any scanner that checks content alone. Does OriginTrace treat a live canonical back-link as declared provenance and de-scope the copy, or does everything unlisted land in the same bucket? That one distinction decides whether syndication keeps working without every mirror having to register first.

Collapse
 
dj29 profile image
Dhruv Jani •

Spot on observation! That exact distinction is why I had to build the "Verdict Engine" on top of the NLP scanner.

If OriginTrace only checked for content overlap, every legitimate cross-post and syndication would get flagged as theft.

To answer your question directly: No, they don't land in the same bucket.

When OriginTrace finds a high-overlap copy, it runs a secondary attribution pass over the DOM. If it finds a live back-link to the original DEV URL (or the original author's name), it flips the hasOriginalLink boolean to true in Sanity. The Verdict Engine then classifies that specific copy as a credited_syndication rather than an unattributed_repost. This automatically de-scopes it from the DMCA generation flow.

Regarding the invisible <link rel="canonical"> tag specifically: currently, the scraper is looking for visible <a> tags in the body to confirm attribution. However, because the extraction pipeline uses Cheerio to parse the full HTML, adding a check for the <head> canonical tag is a brilliant idea and incredibly easy to add to the verdict logic.

The goal is exactly as you said: let syndication flow freely without friction, while catching the bad actors who strip out the links!

BTW thanks for the read! Have A Great Day!

Collapse
 
solo_dev profile image
solo dev •

This one is really good concept, and your future plan with it is also good.

Collapse
 
dj29 profile image
Dhruv Jani •

Thanks!πŸ˜„

Collapse
 
yug_vasava profile image
Yug Vasava •

Good concept and great execution bro, ALL THE BEST!

Collapse
 
dj29 profile image
Dhruv Jani •

Thanks bro!

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The verdict engine is the part I'd want to stress-test hardest, since labeling something unattributed_repost when it was legitimate syndication has real DMCA fallout. Word-overlap plus 5-gram shingling plus LCS is a reasonable ensemble, but I'd want a precision number on the "unattributed" class specifically before letting it act. Did you hold out a labeled set of known-syndicated versus stolen pairs to measure false positives?

Collapse
 
richard_smith_154156d471ef profile image
Richard Smith •

The DMCA draft generation alone makes this worth it. Drafting those from scratch is such a painβ€”having the agent pull the evidence and structure it automatically is a huge time saver.