DEV Community

Dhruv Jani
Dhruv Jani Subscriber

Posted on

OriginTrace: Protecting the DEV Community from Content Theft using Sanity Context MCP

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

A few weeks ago, I was checking my DEV.to dashboard and noticed something weird in the Traffic Sources panel. Views were coming from websites I'd never heard of β€” random blogs, content farms, sites I'd never visited, let alone posted on.

So I clicked through. And there it was: my article, word for word, on someone else's website. Some had credited me. Most hadn't.

That moment of "wait, is this... stolen?" is something every content creator eventually hits. But figuring out whether a repost is actually plagiarism or a legitimate cross-post with credit is surprisingly hard. You have to manually compare the text, check if they mentioned your name, see if they linked back. And if they didn't? Good luck drafting a DMCA notice from scratch.

OriginTrace automates the entire thing.

It's an AI-powered content provenance agent with three core capabilities:

  1. πŸ” The Scanner β€” Paste any DEV.to article URL. OriginTrace extracts distinctive phrases, searches the entire web via Google (Serper API), scrapes up to 15 candidate pages, and runs a multi-signal overlap analysis (word overlap, 5-gram shingling, LCS ratio) to determine if they copied your content.

  2. βš–οΈ The Verdict Engine β€” For every copy found, OriginTrace doesn't just say "it's a match." It checks structured attribution evidence: Does the copy mention the original author's name? Does it link back to the original post? Does it use attribution phrases like "via" or "originally published on"? Based on these boolean signals, it renders a verdict: credited_syndication (legal, no action needed) or unattributed_repost (DMCA time).

  3. πŸ’¬ The Sanity AI Agent β€” This is the heart of the Path One submission. A conversational AI agent powered entirely by the Sanity Knowledge Base and Context MCP. It lets you talk to your provenance data. Ask it: "Which copycat had the highest overlap but still credited me?" or "Draft a DMCA takedown for the worst offender." The agent dynamically queries the structured content in Sanity and gives you actionable, source-backed answers.

Every article, every provenance check, every attribution signal, and every DMCA template is persisted as structured JSON documents in Sanity's Content Lake β€” creating an immutable, queryable ledger of who copied what, when, and whether they gave credit.

Demo

πŸ”— Live App: origintrace.onrender.com

πŸ“Ί Video Walkthrough:

To test the Sanity AI Agent:
(Note: For the best experience, please scan a DEV.to post on the homepage first so the Sanity Knowledge Base has your structured data to query!)
Visit origintrace.onrender.com/chat and try asking:

  • "Which article has the highest overlap percentage?"
  • "Did the top copycat credit the original author?"
  • "Draft a DMCA takedown notice for the worst offender."

Agent querying Sanity Context MCP with transparent reasoning
The AI Agent querying Sanity via Context MCP, with transparent "Agent Thoughts" blocks showing its reasoning process.

Agent delivering a structured attribution analysis
A structured response: the agent analyzed boolean attribution fields from Sanity, rendered a verdict table, and concluded that no DMCA was needed because the copy properly credited the author.

OriginTrace Scanner UI
The scanner homepage β€” paste a DEV.to URL and the programmatic pipeline crawls the web for copies.

Animated radar scanner
Real-time animated radar scanner during an active web crawl.

Scan results with verdicts
Scan results with overlap percentages, attribution verdicts, and DMCA generation β€” all backed by Sanity.

Code

GitHub logo JaniDhruv / OriginTrace

OriginTrace uses Sanity Knowledge Base reconciliation as the provenance engine. Your published articles are the canonical source; a checked URL becomes a KB source; Context's reconciliation surfaces the conflict with citations through MCP.

OriginTrace Logo

OriginTrace

AI-Powered Content Plagiarism Detective & DMCA Agent
Built for Path One: Ship an Agent That Queries Real Content β€” Sanity Challenge 2026

Sanity AI Challenge 2026 Next.js 16 Sanity NVIDIA NIM TypeScript

Live Demo Β· GitHub Β· How It Works Β· The AI Agent


🧠 TL;DR

Paste a DEV.to article URL β†’ OriginTrace's programmatic pipeline crawls the web for stolen copies β†’ structured evidence is persisted to the Sanity Content Lake β†’ then chat with the AI Agent (powered by NVIDIA NIM + Sanity Context MCP) to analyze the data, determine attribution verdicts, generate DMCA takedown notices, and even trigger new scans directly from the chat.

Why this only works with structured content: The agent doesn't keyword-search for answers β€” it queries boolean attribution flags (hasAuthorName, hasOriginalLink), exact overlapPercent integers, and verdict enums from Sanity. A keyword search would never be able to answer "Which copycat had the highest overlap but still credited the…




How I Used Sanity

The Core Thesis: This Agent Only Works Because the Content Is Structured

OriginTrace's AI Agent doesn't do keyword search. It queries typed fields in Sanity β€” booleans, integers, enums, and string arrays β€” to answer questions that would be impossible with unstructured text.

Here's a concrete example. When a user asks:

"Which copycat had the highest overlap but still credited the original author?"

The agent connects to Sanity Context MCP and queries the Content Lake. It needs to:

  1. Filter provenanceCheck documents where verdict != "no_match"
  2. Sort by overlapPercent (a number field) descending
  3. Check attribution.hasAuthorName (a boolean) and attribution.hasOriginalLink (a boolean)
  4. Read attribution.signals (a string array) for the exact credit items found

No keyword search in the world can answer that. The answer requires comparing a number against booleans across multiple documents. That's why structured content matters.

Sanity Schema Design

Four document types power everything:

article β€” The canonical source of truth. When a user scans a DEV.to URL, the pipeline fetches the article, extracts its content, and persists it in Sanity with a reference to its author. This establishes the provenance claim.

provenanceCheck β€” The evidence record. Every candidate copy found on the web gets its own check document containing:

  • overlapPercent (number) β€” similarity score from the multi-signal NLP engine
  • verdict (enum) β€” credited_syndication, unattributed_repost, or no_match
  • attribution (object) β€” structured booleans: hasAuthorName, hasOriginalLink, hasAttributionPhrase, isProperlyAttributed, plus signals[] and missing[] arrays
  • dmcaTemplate (text) β€” a pre-generated legal takedown notice
  • reconciledEntries (array) β€” matched passages between original and copy

author β€” Referenced by articles, contains name and handle for attribution matching.

user β€” Account-level document with email, authId, and avatarUrl for future per-user gating.

Sanity Context MCP Integration

The AI Agent connects to Sanity Context MCP via the AI SDK (@ai-sdk/mcp with HTTP transport). At runtime, the agent:

  1. Calls initial_context to get a Knowledge Base outline
  2. Uses groq_query to dynamically write and execute GROQ queries against the Content Lake
  3. Interprets the structured JSON results (booleans, numbers, arrays) to formulate human-readable answers
  4. Cites the source documents in its responses
// 1. Connect to Sanity Context MCP
const mcpClient = await createMCPClient({
  transport: {
    type: 'http',
    url: SANITY_CONTEXT_ENDPOINT,
    headers: { Authorization: `Bearer ${mcpToken}` },
  },
});

// 2. Fetch MCP tools & combine with our custom scan_article action tool
const mcpTools = await mcpClient.tools();
const tools = { ...mcpTools, scan_article: scanArticleTool };

// 3. Stream response using NVIDIA NIM with injected Sanity tools
const result = streamText({
  model: nim.chatModel('nvidia/nemotron-3.5-lightning-30b-a3b'),
  tools,
  messages,
  system: agentSystemPrompt,
});
Enter fullscreen mode Exit fullscreen mode

Pre-Fetched Structured Context (Not Just MCP)

Before the agent even starts reasoning, the backend runs a dedicated fetchSanityContext() function that executes GROQ queries directly against Sanity to build a structured snapshot of all indexed articles and their provenance checks β€” including per-article breakdowns of actionable vs. credited counts, attribution booleans, overlap percentages, and aggregate platform stats. This pre-fetched context is injected into the agent's system prompt so it can answer summary questions ("How many articles have been scanned?") instantly, without burning an MCP round-trip.

What The Agent Actually Does With Retrieved Content

Sanity Field Type Agent Behavior
overlapPercent number Ranks copycats: "the highest overlap is 98%"
verdict string (enum) Filters results: "3 unattributed reposts found"
attribution.hasAuthorName boolean Evaluates credit: "the copy does mention your name"
attribution.hasOriginalLink boolean Evaluates backlinks: "no link to the original post"
attribution.signals string[] Explains what credit exists: "uses 'via' and links to the original"
attribution.missing string[] Explains gaps: "missing author name and original URL"
dmcaTemplate text Retrieves and customizes the pre-generated DMCA notice

Sanity Project Details

Sanity Project ID: iossngh3
Dataset: production

The schema is defined in TypeScript at sanity/schemas/ β€” the core document is provenanceCheck.ts with its nested attribution object containing the boolean evidence fields.

Agent Session

This project was built iteratively across multiple agent-assisted IDE sessions over the course of the challenge period. The README contains a full technical breakdown of the agent architecture, schema design decisions, and the complete pipeline flow β€” including ASCII architecture diagrams showing how every component connects.

Technical Highlights

⚑ Lightning Fast Agent Reasoning with NVIDIA NIM

Action Agents that use tools (like Sanity Context MCP) require multiple "thinking" loops before they finally respond to the user. If the underlying LLM is slow, the chat experience feels broken.

To solve this, the OriginTrace Agent is powered by NVIDIA NIM, specifically running the Nemotron 3.5 Lightning model. This provides the massive context window needed to ingest complex Sanity JSON payloads, while delivering the ultra-low latency required for real-time tool calling.

🧠 Custom Multi-Signal Overlap Engine

We didn't just want to rely on basic keyword matching. OriginTrace implements a robust NLP engine in TypeScript that resists common plagiarism obfuscation (like deleting a sentence or swapping synonyms). It calculates an overall overlapPercent using a weighted trio of algorithms:

  1. Word Overlap (20%): Jaccard similarity of shared vocabulary.
  2. 5-Gram Shingling (50%): Detects structurally identical paragraphs by matching overlapping 5-word chunks.
  3. Longest Common Subsequence / LCS (30%): Handles reordering and injected words by finding the longest contiguous identical string.

Future Enhancements (Beyond the Hackathon)

OriginTrace is currently a fully functional prototype built in hackathon scope. The AI Agent and Sanity Ledger are currently in Beta, operating as a global, public ledger (similar to a blockchain explorer) where anyone can query all scanned DEV.to URLs.

Here is the future roadmap:

  • πŸ” User Authentication: Add NextAuth/JWT so each author has a private dashboard and personal AI Agent session that only queries their specific provenance data.
  • 🌐 Multi-Platform Support: Extend the scanner to index Hashnode, Medium, and personal blogs.
  • πŸ“¬ Automated Monitoring: Cron-based scheduled scans to proactively alert authors when new stolen copies appear.
  • πŸ€– Paraphrase Detection: Upgrade the NLP engine with LLM embeddings to catch "rewritten" content that evades traditional word-overlap algorithms.

Tech Stack

Component Technology
Frontend Next.js 16 (App Router)
AI Agent LLM NVIDIA NIM (Nemotron 3.5 Lightning)
Agent Framework AI SDK (ai + @ai-sdk/mcp + @ai-sdk/react)
MCP Bridge Sanity Context MCP (HTTP transport)
Knowledge Base Sanity Content Lake (GROQ queries)
Web Search Serper.dev (Google Search API)
NLP Engine Custom TypeScript (word overlap, 5-gram shingling, LCS)
Content Extraction Mozilla Readability + Cheerio + JSDOM
Deployment Render

Thanks for reading! If you've ever seen your own content on a site you didn't post to, OriginTrace was built for you. πŸ›‘οΈ

Top comments (14)

Collapse
 
dj29 profile image
Dhruv Jani • • Edited

OriginTrace is officially live. πŸ›‘οΈ
It does 3 things:
πŸ”Scan the web for copies of your DEV.to posts
βš–οΈ Check whether those copies actually credited you
πŸ’¬ Ask the AI Agent to analyze the evidence and generate a DMCA notice when needed.
So here's the fun part: if you found your own post copied tomorrow, which of these 3 would you use first? πŸ‘€

Collapse
 
dj29 profile image
Dhruv Jani •

Sir @francistrdev, I'd love an opinion in here. Cause after Hacktoberfest gets over, I'm planning to continue this OSS.πŸ˜…

But no rush, give me review when you're free! Have A Great Day!πŸ˜„

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

Solid framing. One gap I keep hitting in practice is proving who accessed training or customer data after the fact. Without signed access logs and a clear purpose tag on each read, takedown and breach response both stall. Curious how you would bind that to an MCP or agent tool boundary. iin1004h2028

Collapse
 
dj29 profile image
Dhruv Jani •

Thanks for the thoughtful question! That is a massive gap in enterprise AI deployments right now.

Currently, OriginTrace is acting as a public ledger, but if we were to lock this down for private customer or training data, binding identity and purpose to the MCP boundary would be critical. Here is exactly how I would architect that integration:

1. Identity Propagation via MCP Context
When the Next.js frontend initializes the AI SDK and connects to the MCP server, we'd pass a signed JWT in the connection headers. The MCP server validates this token before exposing any tools or resources.

2. Tool-Level Purpose Tagging
We would modify the agent's system prompt and the MCP tool definitions to make a purpose argument strictly required. For example, the agent couldn't just call fetch_provenance(). It would have to explicitly call fetch_provenance(articleId, purpose="dmca_investigation").

3. Immutable Audit Logging in Sanity
Since we're already using Sanity Content Lake for structured data, we would use it as the audit trail. Before the MCP server returns the requested data to the agent, it writes an accessLog document to Sanity containing:

  • The timestamp
  • The userId (extracted from the JWT)
  • The action (the specific MCP tool called)
  • The purpose (provided by the agent)

By intercepting the request at the MCP server level, we guarantee that the LLM cannot bypass the logging mechanism. The Sanity Content Lake becomes a fully queryable, timestamped ledger of every agent action for compliance teams.

It’s definitely the next logical step for secure agentic systems! Thanks again for reading.

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

The public-ledger angle is the interesting one. We syndicate on purpose β€” mirrors carry canonical_url back to the original β€” and the only thing separating that from theft is the declared link, which is invisible to any scanner that checks content alone. Does OriginTrace treat a live canonical back-link as declared provenance and de-scope the copy, or does everything unlisted land in the same bucket? That one distinction decides whether syndication keeps working without every mirror having to register first.

Collapse
 
dj29 profile image
Dhruv Jani •

Spot on observation! That exact distinction is why I had to build the "Verdict Engine" on top of the NLP scanner.

If OriginTrace only checked for content overlap, every legitimate cross-post and syndication would get flagged as theft.

To answer your question directly: No, they don't land in the same bucket.

When OriginTrace finds a high-overlap copy, it runs a secondary attribution pass over the DOM. If it finds a live back-link to the original DEV URL (or the original author's name), it flips the hasOriginalLink boolean to true in Sanity. The Verdict Engine then classifies that specific copy as a credited_syndication rather than an unattributed_repost. This automatically de-scopes it from the DMCA generation flow.

Regarding the invisible <link rel="canonical"> tag specifically: currently, the scraper is looking for visible <a> tags in the body to confirm attribution. However, because the extraction pipeline uses Cheerio to parse the full HTML, adding a check for the <head> canonical tag is a brilliant idea and incredibly easy to add to the verdict logic.

The goal is exactly as you said: let syndication flow freely without friction, while catching the bad actors who strip out the links!

BTW thanks for the read! Have A Great Day!

Collapse
 
solo_dev profile image
solo dev •

This one is really good concept, and your future plan with it is also good.

Collapse
 
dj29 profile image
Dhruv Jani •

Thanks!πŸ˜„

Collapse
 
yug_vasava profile image
Yug Vasava •

Good concept and great execution bro, ALL THE BEST!

Collapse
 
dj29 profile image
Dhruv Jani •

Thanks bro!

Collapse
 
jays_tech profile image
Jay •

The split between credited_syndication and unattributed_repost is the part most plagiarism scanners skip, and leaning on boolean attribution signals (name mention, backlink, "via") keeps it from nuking legitimate reposts. My worry would be the 5-gram shingling on heavily paraphrased theft that keeps the ideas but swaps the words. Have you seen the LCS ratio hold up there, or does that case slip through?

Collapse
 
richard_smith_154156d471ef profile image
Richard Smith •

The DMCA draft generation alone makes this worth it. Drafting those from scratch is such a painβ€”having the agent pull the evidence and structure it automatically is a huge time saver.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.