What if an AI system could take a pile of unstructured claim documents and turn them into a traceable map of claims, evidence, and events?
That was the idea behind ClaimPilot, our AI-powered evidence intelligence system built for the Hacktoberfest Hack Day Coimbatore × INIT Club & IDEA Club.
Instead of simply asking an LLM to summarize a document, ClaimPilot focuses on something more useful for investigation: connecting pieces of evidence and preserving where they came from.
🚀 What We Built
ClaimPilot takes unstructured claim-related documents and transforms them into structured case intelligence.
The pipeline looks like this:
Claim Documents
↓
Document Processing
↓
Gemma 4 31B IT
↓
Claims + Evidence + Events
↓
Evidence Graph
↓
Timeline
↓
Hugging Face Embeddings
↓
Pinecone
↓
Semantic Evidence Retrieval
↓
Contradiction Analysis
The goal is to make a complex case easier to investigate by connecting:
Documents → Claims → Evidence → Events
For example, instead of just producing a summary such as:
"The device stopped working and was inspected."
ClaimPilot can represent the information as structured intelligence:
Claim:
Device stopped functioning
Evidence:
User reported device failure
Event:
Consumer reported failure
Date: August 2
Evidence:
Internal diagnostics indicated a power surge
Event:
Power surge detected
Date: August 2, 14:30 UTC
Claim:
Device may require replacement
Evidence:
Inspection report recommended power board replacement
This structure makes the information much more useful for downstream reasoning.
🧠 Why Gemma 4?
The core intelligence layer of ClaimPilot uses Gemma 4 31B IT.
We use Gemma to process the document and extract three important types of information:
1. Claims
Claims are substantive assertions made in the document.
For example:
"The device stopped functioning."
"No external physical damage was found."
"The internal diagnostics indicate a power surge."
We deliberately distinguish these from document metadata such as document IDs or document dates.
2. Evidence
Evidence represents information that can support a claim.
For example:
Evidence:
Internal diagnostic logging indicates a power surge.
Supports:
Claim → Device failure may have been caused by an electrical event.
3. Events
Events represent things that happened at a particular point in time.
For example:
2026-08-02
Consumer reported that the device stopped working.
2026-08-05
Technician inspected the device.
2026-08-02 14:30 UTC
Power surge was detected.
This gives us both a semantic view of the case and a chronological view.
🔗 Building the Evidence Graph
One of the most important parts of ClaimPilot is the evidence graph.
We don't want the AI output to become a black-box paragraph that nobody can trace.
Instead, we represent relationships explicitly:
Document
│
├── contains → Claim
│ │
│ └── supported by → Evidence
│
├── contains → Evidence
│ │
│ └── associated with → Event
│
└── contains → Event
For example:
Claim_001
↓ supported by
Evidence_001
↓ associated with
Event_001
This allows us to answer questions such as:
- What evidence supports this claim?
- Which document did the evidence come from?
- Which events are related to the claim?
- What happened before or after the event?
- Which pieces of evidence should be compared?
This traceability is one of the main design principles of ClaimPilot.
⏱️ Building a Case Timeline
The same extracted events are used to construct a chronological timeline.
Instead of having important dates buried inside multiple documents, ClaimPilot organizes them into a single sequence.
Aug 02
│
├── Consumer reports device failure
│
└── 14:30 UTC
Power surge detected
Aug 05
│
└── Technician inspection
Aug 07
│
└── Inspection report filed
Replacement recommended
We also preserve precise timestamps when they are available instead of unnecessarily reducing everything to a date.
This becomes important when the order of events affects the interpretation of evidence.
🔎 Semantic Evidence Retrieval
Another layer of ClaimPilot uses embeddings.
We use a Hugging Face-hosted embedding model to convert evidence and document text into vectors.
The architecture is:
ClaimPilot Backend
↓
Hugging Face Inference API
↓
1024-dimensional embedding
↓
Pinecone
The embedding model runs through Hugging Face's hosted inference rather than being downloaded and executed locally.
The resulting vectors are stored in Pinecone.
This gives ClaimPilot semantic retrieval capabilities.
For example, if the system needs evidence related to:
"electrical damage"
it can retrieve semantically related evidence even when the document uses different wording such as:
"power surge"
"voltage event"
"electrical fault"
This is particularly useful when investigating evidence across multiple documents.
⚙️ How We Built It
We built ClaimPilot as a modular backend pipeline.
The major services are separated by responsibility:
document_service.py
↓
gemma_service.py
↓
vector_service.py
↓
evidence_service.py
↓
timeline_service.py
Document Service
Responsible for extracting usable text from uploaded documents.
Gemma Service
Responsible for structured extraction:
Document
↓
Gemma 4
↓
Claims
Evidence
Events
Vector Service
Responsible for:
- generating embeddings through Hugging Face
- storing document embeddings
- storing evidence embeddings
- querying Pinecone for semantic retrieval
Evidence Service
Responsible for constructing meaningful relationships between:
- documents
- claims
- evidence
- events
Timeline Service
Responsible for:
- ordering events chronologically
- preserving timestamps
- associating events with their source information
Keeping these responsibilities separate makes it easier to extend the system with additional reasoning capabilities.
🛠️ Development with Antigravity
We built the project iteratively using Antigravity as our AI-assisted development environment.
Instead of trying to generate the entire application in one step, we broke the system into smaller services and milestones.
The development process was roughly:
1. Define the evidence intelligence architecture
↓
2. Build document processing
↓
3. Integrate Gemma 4
↓
4. Implement structured extraction
↓
5. Build embeddings pipeline
↓
6. Integrate Pinecone
↓
7. Build evidence graph
↓
8. Build chronological timeline
↓
9. Validate the complete pipeline
↓
10. Prepare contradiction reasoning
One important lesson from this process was that integrating AI models is only one part of building an AI application.
The difficult part is making sure that the output is:
- structured
- traceable
- retrievable
- consistent
- useful for downstream reasoning
That's why ClaimPilot isn't designed as simply:
Document → LLM → Summary
Instead, we're building:
Document
↓
Structured Intelligence
↓
Relationships
↓
Semantic Retrieval
↓
Reasoning
🔮 What's Next?
The next major component we're working toward is the Contradiction Engine.
The idea is to combine the evidence graph with semantic retrieval.
For a particular claim, the system can retrieve relevant evidence and compare the evidence across documents.
Conceptually:
Claim
↓
Find supporting evidence
↓
Find related evidence
↓
Compare evidence
↓
Identify:
├── Supporting
├── Contradicting
└── Uncertain
This could allow ClaimPilot to identify situations such as:
Document A:
"No physical damage was observed."
Document B:
"External damage was visible on the device."
↓
Potential contradiction
The goal is not simply to make the AI produce an answer, but to make the reasoning traceable back to the underlying evidence.
🧩 Technology Stack
AI / Models
- Gemma 4 31B IT
- Google Gemini API
- Hugging Face hosted inference
Data / Retrieval
- Pinecone
- Vector embeddings
- Evidence graph
- Timeline construction
Backend
- Python
- FastAPI
- Modular service architecture
Development
- Antigravity
- GitHub
🎯 Why We Built ClaimPilot
The larger idea behind ClaimPilot is simple:
AI should not just tell you what a document says. It should help you understand how the pieces of evidence connect.
Large document-heavy investigations contain claims, statements, dates, reports, and supporting evidence scattered across many sources.
ClaimPilot tries to turn that unstructured information into something investigators can actually reason over.
From:
Hundreds of pages of documents
to:
Claims
+
Evidence
+
Events
+
Relationships
+
Semantic Retrieval
And eventually:
Traceable AI-assisted investigation
🔗 Project Links
GitHub: [https://github.com/djivites/HTF044_ZERO_LATENCY_CLAIM_PILOT]
Demo Video: [https://github.com/djivites/HTF044_ZERO_LATENCY_CLAIM_PILOT]
If you're interested in AI systems that combine LLMs + structured data + retrieval + reasoning, we'd love to hear your thoughts.
Top comments (0)