This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built this for my friend — a freelance developer who routinely puts in 10-to-12-hour days across multiple machines, three monitors, and a chaos of browser tabs, Docker containers, and half-finished code sessions. His single biggest frustration is what he calls "the Monday morning blank screen": he sits down, opens his laptop, and has absolutely no idea what he was doing when he stopped. He loses 30-to-45 minutes every single morning just reconstructing context — what file was he editing, what Docker error was he chasing, what did he promise to send to a client before he closed the laptop.
TigerDB is a unified desktop memory intelligence system that watches his work sessions, stores every meaningful activity into a three-layer database, and hands him a concise morning briefing the moment he opens his machine.
The three layers are:
-
Vector Engine (
pgvector+ HNSW index): encodes every file edit, terminal command, and browser visit as a 384-dimensional embedding for natural-language search -
Property Graph Engine (
tiger_graph): models each work session as a directed acyclic graph — activity nodes connected by typed edges (ACCESSED_RESOURCE,COMMITTED_TO) — so you can trace exactly what led to what - TimescaleDB Hypertable: a time-partitioned telemetry log that tracks which workflows succeeded, which failed, and the GRPO reward signals that teach the agent to get smarter session over session
When my friend opens his laptop in the morning, a single click surfaces three tiles: yesterday's unfinished threads with their progress percentage, any blockers or failed deployments that need Reflect-Replan, and today's auto-generated agenda. He can also speak a natural query — "where did I put that Kafka config?" — and the agent retrieves the exact file path, code snippet, and the promise he made about it, ranked by a hybrid scoring formula that weights semantic similarity, past success rate, and recency.
Demo
Live Site: TigerDB: The Desktop Memory Agent
The demo shows:
- TigerDB's memory architecture
- Property Graph execution trajectories
- Live desktop activity insertion
- Vector memory retrieval
- Equation 4 memory ranking
- Positive and Negative Paradigms
- TimescaleDB telemetry
- Voice-based memory retrieval
Running Locally
docker compose up
npm run dev
uvicorn tiger_agent.api:app --host 0.0.0.0 --port 8000
The application runs locally with PostgreSQL, pgvector, TimescaleDB, React, and FastAPI.
Code
TigerDB: Unified Graph & Vector Database for Memory Intelligence Agent (MIA)
An enterprise-grade implementation of the Memory Intelligence Agent (MIA) lifelong learning framework (arXiv:2604.04503v4), powered by TigerDB — a unified multi-model engine combining PostgreSQL 18, TimescaleDB HA, pgvector, and a native Property Graph Model.
Includes a full-stack Desktop Memory Agent Web HUD with real-time Graph DAG visualization, speech-enabled AI assistant, morning briefings, in-situ WebMCP auto-healing, and a zero-downtime deployment container for Render.com.
📑 Table of Contents
-
1. Theoretical Foundations (MIA Paper arXiv:2604.04503)
- 1.1 Equation 4: Non-Parametric Hybrid Retrieval Scoring
- 1.2 Multimodal Semantic Similarity (Eq. 9)
- 1.3 Dual Paradigm Extraction (Positive vs. Negative Paradigms)
- 1.4 Shortest Path Execution Prioritization
- 1.5 High Semantic Similarity Knowledge Replacement
- 1.6 Reflect-Replan Mechanism
- 1.7 GRPO Trajectory Telemetry in TimescaleDB
- 2. System Architecture
- 3. Application Features & Web HUD
- 4.…
Key files:
-
tiger_agent/desktop_memory.py— core activity recorder: creates graph nodes + vector embeddings in a single transaction -
tiger_agent/morning_assistant.py— morning briefing engine: unfinished threads, blockers, today's agenda -
tiger_agent/llm_client.py— NVIDIA NIM gateway with structured JSON output and local fallback -
tiger_mia/— Memory Intelligence Agent core: Manager, Planner, Executor, Judger
How I Built It
The architecture implements the Memory Intelligence Agent (MIA) framework (arXiv:2604.04503) on top of PostgreSQL 18. Every component is open-source:
Inference layer — NVIDIA NIM (open-weight models via NIM API):
The agent uses meta/llama-3.3-70b-instruct served through NVIDIA NIM — an open-weight model that can be self-hosted on any NVIDIA GPU. The LLM client (llm_client.py) uses langchain_openai pointed at the NIM endpoint and requests structured JSON output via Pydantic schemas, so the response is always parseable into { format, speakingtext, content } without hallucinated formatting.
Embedding layer — sentence-transformers (local inference):
Every desktop activity is encoded into a 384-dimensional vector using all-MiniLM-L6-v2, running fully locally via the sentence-transformers library. No activity text ever leaves the machine at embedding time.
Memory layer — TigerDB (PostgreSQL 18 + pgvector + TimescaleDB):
-
pgvector 0.8.6provides HNSW indexing (vector_cosine_ops) for sub-millisecond cosine similarity search - A custom
tiger_graphschema stores execution DAGs in plain PostgreSQL tables, traversed with recursive CTEs - TimescaleDB hypertable partitions the trajectory log into 7-day chunks with 90% columnar compression
Agent runtime — four-role architecture:
| Role | Responsibility |
|---|---|
| Memory Manager | Hybrid retrieval (Eq. 4), knowledge replacement at threshold θ=0.92, workflow compression |
| Planner | Few-shot in-context learning with Positive & Negative Paradigms, Reflect-Replan |
| Executor | ReAct loop tool interactions, logging execution graphs into TigerDB |
| Judger | Evaluates correctness, computes multi-signal GRPO rewards (r₁ correctness, r₂ tool efficiency, r₃ format) |
The hybrid scoring formula (Equation 4 from the MIA paper) ranks retrieved memories by:
This means a memory that was correct and short rises above memories that were merely semantically close — the agent genuinely learns which workflows worked.
🐯 Tiger Data: Best Use of Tiger Data
TigerDB was designed from the ground up around exactly the three patterns the Tiger Data prize calls out: store embeddings with pgvector, run hybrid keyword and vector search for an agent, and let an agent manage a database. Here is precisely how each one works.
1. Storing Embeddings with pgvector
Every desktop activity — a file edit, a terminal command, a browser tab, a promise made to a client — is encoded into a 384-dimensional vector using sentence-transformers/all-MiniLM-L6-v2 and persisted into tiger_mia.memory_units via pgvector:
# tiger_agent/desktop_memory.py — record_activity()
q_embed = self.embedder.encode(
f"{title} {activity_type} {path_or_url} {snippet} {promise_text or ''}"
)
insert_sql = """
INSERT INTO tiger_mia.memory_units (
modality, category, question, caption,
question_embedding, -- vector(384)
trajectory_graph_id,
compressed_workflow, judgment_label,
execution_length, usage_count, success_count
) VALUES (
'desktop_activity', %s, %s, %s,
%s::vector, %s,
%s::jsonb, %s,
1, 1, 1
) RETURNING id;
"""
The HNSW index (vector_cosine_ops) makes every lookup sub-millisecond even as memories accumulate across months of work sessions.
2. Hybrid Keyword + Vector Search for an Agent
The search_timeline method — called every time the agent answers a voice query — runs a true hybrid scoring formula directly in SQL, combining pgvector cosine similarity with PostgreSQL full-text similarity() in a single query:
-- tiger_agent/desktop_memory.py — search_timeline()
SELECT
id,
question AS title,
caption AS snippet,
category AS activity_type,
compressed_workflow,
trajectory_graph_id,
-- Hybrid score: 70% keyword (pg_trgm similarity) + 30% vector (pgvector cosine)
(0.7 * GREATEST(similarity(question, %s), similarity(COALESCE(caption, ''), %s))
+ 0.3 * (1 - (question_embedding <=> %s::vector))) AS composite_score,
(1 - (question_embedding <=> %s::vector)) AS similarity
FROM tiger_mia.memory_units
WHERE modality = 'desktop_activity'
ORDER BY composite_score DESC
LIMIT %s;
When the Planner retrieves memories for in-context learning, it uses Equation 4 from the MIA paper — a three-signal scoring formula that weighs semantic similarity, historical success rate, and access frequency together:
This means the agent doesn't just retrieve what's similar — it retrieves what worked, actively separating Positive Paradigms (correct past workflows) from Negative Paradigms (failed patterns to avoid).
3. Agent Managing the Database Through WebMCP
TigerDB includes a WebMCP Action Engine — a direct agent-to-database execution layer that lets the agent manage system state without the user copy-pasting commands. When the morning briefing surfaces a failed Docker deployment (stored as a Negative Paradigm in TigerDB), the agent calls docker_heal directly:
# tiger_agent/webmcp_actions.py
class WebMCPActionEngine:
def execute(self, action_type: str, params: Dict[str, Any]) -> Dict[str, Any]:
"""Dispatches action to the appropriate handler."""
handler_map = {
"docker_heal": self.heal_docker_container, # auto-restarts containers
"docker_status": self.get_docker_status, # live container table
"open_file": self.open_local_file, # reveals file in explorer
"draft_email": self.create_draft_email, # creates draft from promise
"resume_session": self.resume_session # restores work contexts
}
And on the database side, the agent writes every action outcome — whether the Docker heal succeeded or failed — back into tiger_mia.memory_units as a new judgment-labeled memory, and logs the reward signals to the TimescaleDB hypertable:
-- TimescaleDB Hypertable: agent self-updates its own memory after each action
INSERT INTO tiger_mia.trajectory_logs
(time, question, total_steps, correctness_reward, tool_reward, format_reward,
total_reward, advantage, metadata)
VALUES (%s, %s, %s, %s, %s, %s, %s, %s, %s::jsonb);
The result is a closed loop: the agent reads from TigerDB, acts on the world, and writes the outcome back to TigerDB — with no human in the middle.
Why Does Open Innovation Matter?
My friend's computer is his livelihood. Every file path, every client name in a browser tab, every half-written email draft is private business data. The core insight of this project is that desktop memory is the most personal data a person generates, and it should never leave their machine by default.
With closed AI APIs, continuous desktop activity monitoring would require streaming every file name, every snippet of code, every URL visited to a third-party server for embedding and inference. That is a non-starter for a freelancer with NDAs and client confidentiality obligations.
Open-source AI made three things possible that a closed API never could:
Local embedding inference.
sentence-transformerswithall-MiniLM-L6-v2runs entirely on the CPU. Every activity is vectorized locally. The data never moves off the machine.Self-hostable LLM. NVIDIA NIM serves open-weight models like Llama 3.3 70B. My friend can run this on his own NVIDIA GPU at home. The morning briefing LLM call — the part that reads his unfinished tasks and failed deployments — stays on his own hardware.
Inspectable, modifiable memory. Because TigerDB is built on PostgreSQL with documented schemas, my friend can open
psql, see exactly what the agent remembers about him, delete rows he doesn't want, and modify the scoring weights. A closed memory API gives you none of that transparency or control.
Open innovation didn't just make TigerDB cheaper — it made it ethically possible to build a tool that watches someone's work all day.
My Agent Session
Prize Categories
- 🐯 Tiger Data — Best Use of Tiger Data: pgvector HNSW embeddings for all desktop activities + hybrid keyword/vector SQL search + WebMCP agent that reads, acts, and writes outcomes back to TigerDB autonomously
- 🟢 NVIDIA: Uses NVIDIA NIM API with open-weight Llama 3.3 70B for structured LLM responses
- 🐘 TimescaleDB / PostgreSQL: Core storage layer — pgvector HNSW + TimescaleDB hypertable + recursive CTE property graph, all on PostgreSQL 18
- 🦜 LangChain:
langchain_openaiused as the LLM client interface to NVIDIA NIM
Built during Hacktoberfest Weekend 2026.

Top comments (0)