DEV Community

Cover image for FieldIssue: take a walk, report a problem, come back with evidence
Himanshu Kumar
Himanshu Kumar Subscriber

Posted on AI-assisted

FieldIssue: take a walk, report a problem, come back with evidence

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass.

What I Built

Most "report a pothole" apps stop at the photo. FieldIssue is built around the second walk.

Snap a broken bench, an overflowing bin or a cracked pavement, add a short note, and an open-weight Gemma model writes up what it sees. Walk past again later, take another photo, and FieldIssue compares the two: which conditions were added, which were removed, which haven't changed. Over time, a report becomes a dated trail of evidence instead of a single complaint.

FieldIssue report FI-000007: a real photo of a cut tree with dried branches and plastic litter, which Gemma classified as Environment, Medium

A real walk: I photographed this cut tree and litter in Lucknow on 9 October and reported it from my phone. Gemma listed three conditions: cut or broken branches, piled dried branches and leaves, and plastic waste.

On the same walk I reported a second spot about 50 metres away, FI-000006, where fallen branches block a dirt path. Both reports carry my phone's GPS and capture time. Next I go back to the same spot, from the same angle, and let FieldIssue compare.

FieldIssue report FI-000006: a real photo of fallen branches across a dirt path, with Gemma's conditions and the issue history

Two rules keep it honest:

  • The model recommends, a person decides. Gemma can suggest a status, but only a human resolves a report, and a manual resolution is labelled as one in the history.
  • Identical file, no new evidence. Upload the same image twice and you get "Repeated photo: no new evidence." It's an exact-byte check: it stops lazy re-uploads, not a determined faker.

Safeguard test on the hosted demo: the same sample file uploaded twice, so the saved comparison flags a repeated photo as no new evidence

The safeguard in action, tested with a sample photo and placeholder coordinates.

Beyond single reports, you can plan a walk past nearby issues, keep private saved walks with an optional account, and run a private community that assigns reports and needs fresh compared evidence plus two member approvals to close one.

Demo

Open FieldIssue · Explore reports · Model Lab

A 2-minute tour (no sign-up needed; the free Render instance may take a few seconds to wake):

  1. Open a real report from my walk, then a sample report that shows the repeated-photo safeguard.
  2. Play its ElevenLabs audio briefing, then open the Lab's recorded Tinker examples: all 18 held-out notes with the base and fine-tuned answers side by side, mistakes included.
  3. Create your own report, come back with a new photo, and compare.
  4. In Walk, pick a starting point, add nearby open issues and open the route in Google Maps.

Owner view of the hosted sample report showing its photo, map, audio briefing, repeated-photo warning and saved Tinker result

It's an installable web app: if you lose signal mid-walk, captures are queued on your phone and uploaded only after you review them.

Code

GitHub repository · MIT license · Setup and demo guide

Started on 6 October 2026. A labelled fixture mode runs locally with no API keys; the hosted demo uses the real integrations.

How I Built It

React, Vite, Tailwind, shadcn/ui and Leaflet on the front. A TypeScript Hono API runs typed Mastra workflows for observations and revisits. One free Render container runs the Node API and a private FastAPI process, with Render PostgreSQL/PostGIS as the source of truth. Gemma runs through Google's hosted API, and Tiger Data keeps a separate pgvector index for semantic search.

Issue history on the hosted demo: created from a photo, classified from validated model evidence, title edited by a person, then a manual resolution

FieldIssue architecture: a React workspace with an offline capture queue talks to a Hono API with Mastra workflows, which calls Render PostgreSQL and PostGIS, FastAPI with Gemma, Tiger Data vector search, Tinker, TabPFN and Backboard, SerpApi and ElevenLabs, and Sentry

Each model has one job: Gemma reads photos, Tinker reads the reporter's words, Backboard compares two open models side by side. None of them can close an issue on its own. Model output is schema-validated before it's saved, and offline retries reuse an idempotency key so a dropped connection never creates a duplicate report.

What the fine-tune actually bought

With Tinker I fine-tuned Qwen3-8B on 36 synthetic field notes, including informal Hindi and Hinglish ones. On 18 held-out notes:

Check Base Fine-tuned
Valid schema 18/18 18/18
Correct category 17/18 17/18
Correct severity 13/18 17/18

Severity is the field that matters for triage, and it improved in both training runs: +4 here, +2 in an earlier run on 7 October. Eighteen synthetic notes is a signal, not a benchmark, so the evaluation is public and every answer is browsable in the Lab. The checkpoint serves the live note action on reports you own.

I also tried TabPFN for ranking which issues to walk past. On 24 held-out synthetic rows it scored 16/24 against a 17/24 majority baseline, so I kept it in the Lab and out of real recommendations. Shipping a model that loses to "always guess the common answer" would have been the easy mistake.

FieldIssue Model Lab showing a saved synthetic TabPFN scenario result with data-use consent

What broke

  • Saved walks failed in the browser even though the API tests passed: the location picker was sending display metadata to an endpoint that expected bare coordinates. Lesson: test the whole form, not just a clean payload.
  • My ElevenLabs key expired mid-week, and the cache hid it. Saved briefings kept playing; only a fresh generation returned 401. I now check fresh calls separately from cached ones.
  • Sentry separated a soft failure from a hard one. A test report at a spot with no nearby places got nothing back from SerpApi, but the Gemma report still saved, and Sentry logged it as PLACE_CONTEXT_UNAVAILABLE without leaking notes or photos (event, training trace).

The latest build passes 242 automated tests, plus hosted checks for account isolation, saved walks, membership revocation and audio playback (acceptance ledger).

Why Does Open Innovation Matter?

If an app is going to tell a city "this got worse", people should be able to check how it decided. Anyone can read FieldIssue's prompts, check the output schema, rerun the evaluations and see exactly where the models failed. Open weights also mean the vision model is swappable, and a narrow task like note reading can be fine-tuned and measured against its base. The app is MIT-licensed and was built on free tiers and sponsor credits.

Limits

  • Every model evaluation here uses synthetic data. Next: a human-labelled change history from real revisits.
  • Free tiers run out: the Render database expires on 7 November 2026, the Tinker checkpoint and ElevenLabs key on 8 November. New inference shares a cap of 60 provider units a day; cached results stay viewable.
  • Sentry traces are sampled, so not every request has a complete trace.

My Agent Session

Codex built the backend and ran integration and browser checks. Claude Opus built the frontend and reviewed the later changes. The Entire CLI captured the frontend session; a reviewed excerpt of nine real messages is public. The landing illustration is AI-generated and labelled as such. This write-up was prepared with AI assistance.

Prize Categories

  • Best Use of Gemma: reads every photo and compares each revisit with the original, listing what was removed, added and unchanged.
  • Best Use of Tinker: Qwen3-8B fine-tuned on English, Hindi and Hinglish field notes; severity accuracy rose from 13/18 to 17/18, live in the app.
  • Best Use of Render: one container runs the web app, Node API and private FastAPI process, with Render PostgreSQL/PostGIS as the database.
  • Best Use of TabPFN: scores synthetic revisit scenarios in the Lab; at 16/24 against a 17/24 baseline, I published the result and kept it out of recommendations.
  • Best Use of Mastra: typed workflows run every new observation and every revisit comparison.
  • Best Use of Sentry Agent Tracing: sanitized spans on each workflow step, with soft provider failures like PLACE_CONTEXT_UNAVAILABLE recorded.
  • Best Use of SerpApi: nearby place context for each report; a report still saves when no place is found.
  • Best Use of ElevenLabs: cached spoken briefings for each report.
  • Best Use of Tiger Data: hybrid keyword and pgvector search over reports (all-MiniLM-L6-v2), with a retry queue so indexing never blocks a report.
  • Best Use of Backboard: Gemma 3 and Qwen 2.5 read the same report side by side, so a person can see where they disagree.
  • Best Use of Entire: captured the frontend agent session; a reviewed excerpt is linked above.

Top comments (0)