This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
I hunt bounties and hackathons from my desk in Vietnam, all day. The only time I go outside is for lunch: about 200 metres to buy rice and water, then straight back. I hurry because every minute away feels like a listing I missed.
So I built the thing that would let me walk slower. It is called Đi đây, which is what you say in Vietnamese when you get up to leave: "I'm off."
- When I get up for lunch, I press Đi đây on my computer and scan a QR code with my phone.
- While I am out, Gemma 4, running on my own computer, reads the newest bounty and hackathon listings and checks the fine print. It looks for the catch (a live interview, being there in person, adding a payment card), whether Vietnam is allowed, and the real deadline in Vietnam time.
- The results stay locked until I come back with one photo and ten seconds of sound from outside. Gemma listens to the recording, looks at the photo, decides whether I really went out, and only then opens the list.
It is for people like me: anyone whose work lives in a browser tab and who stays inside because of it. The job it does while you are out can be swapped. Mine is reading listings.
My first idea for this week was different. It was a walking game where the model wrote the rules and made a little zine from your photos afterwards. When I looked at it again, it looked like something a machine would make for a person it had never met. I threw it away and built what I needed.
Demo
There is no hosted link, on purpose: everything runs on my computer at home and the phone only talks to it over home wifi. So here are the first two real trips instead, both on Thursday 8 October.
Lunch, 12:06 to 13:04.
This is where I stopped. On one side is a canal that brings seawater in for a salt company. On the other side are the fields they flood with seawater to make salt. Trucks pass on the road next to it.
Gemma heard: "The sound of a car driving by and the sound of wind." That is right.
Gemma saw: "A wide river flows past a rocky embankment with mountains in the background under a cloudy sky." Close, but it is a canal and salt fields, not a river. It said "outdoors" and opened the list, which was the only thing it had to get right.
What happened while I was out:
- It read 20 listings in 10 and a half minutes, between 23 and 45 seconds each.
- Then it waited for me for another 47 minutes. I had told myself the slow laptop would buy me the walk. It turns out lunch is slower than the laptop.
- It marked 1 listing worth doing, 10 maybe, 9 no.
The "no" list is where it earned its lunch break. It caught two hackathons that start on site in Kraków and Warsaw, a track that asks teams to pick Pakistan as their country, one that is limited to Ukraine, and a demo day where you show the product live. Those are the lines I usually find on the third read, after I already started.
It also got things wrong, and I want to show those too:
- Three listings said "Record a pitch video and a demo video." Gemma called that a live interview, three times. A recorded video is not live. I had already written that into the prompt, and it still did it.
- For one listing it said I would need a payment card. The quote it gave for that never mentions a card. It guessed.
The fix is a rule in code, not a longer prompt: a catch only counts if the words Gemma quotes actually say it. "Live interview" needs a word like live, interview or call in the quote. "Card required" needs card, credit or billing. If the words are not there, the catch is dropped and the list shows that it was dropped. I replayed this trip through the new check: the four wrong catches drop, the four real ones stay, and the Ukraine-only listing is still a no because of its region.
All 20 of those came from Superteam Earn. The queue reads the soonest deadlines first, and at lunch those were all on Superteam.
Same day, 15:18 to 15:33. A short one.
I went out again in the afternoon, for 15 minutes this time. Gemma heard "a low rumble of a vehicle and some indistinct voices" (voices, described and not written down, which is how it should be) and saw "various buildings under a bright, partly cloudy sky."
This time the laptop was the slow one. The new listings were mostly Devpost hackathons with long rules pages, and each took between 43 seconds and almost two minutes. It had read 14 of 19 when I got back, and finished the last five while I was already at my desk. So the walk is not always longer than the work. I like that the list does not care.
What it caught:
- A hackathon whose rules list Vietnam by name among the countries that cannot enter. Gemma quoted the line. I would have found it after registering.
- Three that are only for students, and one only for members of a separate developer programme.
What it got wrong:
- It said no to one hackathon because the page told it to "upgrade your browser to Internet Explorer 10 or higher." That line is a banner Devpost hides in an HTML comment for very old browsers. My code passed it to Gemma as if it were part of the rules. That was my bug, not Gemma's, and it is fixed: comments are stripped before the model reads a page.
- For one hackathon it wrote that Vietnam is not on the excluded list, and then marked Vietnam as not allowed anyway. No rule in code catches that yet. I still read the "no" list myself.
And one thing that made me laugh: three of its "maybe" hackathons are ones I had already entered this month. Gemma was not wrong about them.
Code
ShenJun93
/
di-day
Đi đây ("I'm off"): my laptop reads bounty fine print with Gemma 4 while I walk to buy lunch. Results unlock when I come back with a photo and 10 s of street sound.
Đi đây
Đi đây is Vietnamese for "I'm off". I sit at my computer all day hunting bounties and hackathons. The only time I go outside is a 200-metre walk to buy lunch, and I hurry back because every minute away feels like a missed listing.
So the laptop works while I'm out. When I get up, I press Đi đây. Gemma 4, running on this computer reads the new listings and checks the fine print the way my Bounty Triage benchmark does: hidden gates (a live interview, a card, in-person attendance), whether Vietnam is eligible, and the real deadline in Vietnam time On my old laptop that takes about half a minute per listing. On the first real trip it read 20 listings in ten and a half minutes; lunch took 58.
The results stay locked until I come back with one photo and ten seconds of sound…
Node with no dependencies, about 600 lines. server.mjs runs the trip and the lock, brain.mjs talks to Gemma, sources/ fetches the listings, and public/out.html is the page on the phone.
How I Built It
Model: Gemma 4 E4B, the Q4_0 GGUF, served by llama.cpp's llama-server in Docker. My GPU is a 4 GB Quadro P1000, so only 8 layers go on the GPU and the rest run on the CPU. It is slow, and that is fine here.
One model, three jobs. The same Gemma reads the listing text, listens to a WAV file, and looks at a JPEG, all through llama.cpp's OpenAI-style API.
The phone part. The phone page is served by my computer over home wifi. The browser will not give a plain http:// page on the local network the microphone, so the page uses the phone's own camera and recorder (<input type="file" capture>). Then Web Audio turns the recording into 16 kHz mono WAV before sending it. Outside, the phone needs no signal at all.
Things I learned the hard way, and kept in the code:
- Listen before you look. When I sent the photo and the sound in one prompt, Gemma described traffic in a recording of a speech. It was reading the photo. Now the sound goes in alone first, and the photo second.
- Speech is described, never written down. The recordings have my neighbours in them.
- Time zones are done in code. I made a Kaggle benchmark on bounty fine print a few days earlier, and small models kept getting "11:59 PM PDT" wrong in Vietnam time. So Gemma only copies the deadline and the zone as written, and the code does the maths.
-
Plain JSON mode, not a JSON schema. With a full schema, this model padded its answers with whitespace until it ran out of tokens.
response_format: {type: "json_object"}and a short example in the prompt work better. - A catch needs its words. The rule from the Demo section above.
- Read what the model reads. The Internet Explorer banner was in the text I sent, not in the model. When an answer looks strange, I check the input first now.
Why Does Open Innovation Matter?
Because of what goes into it. Every trip sends ten seconds of my street and a photo of where I stand, and my whole list of things I am trying to win. With an open-weight model, none of that leaves my desk. I would not send recordings of my neighbours to a server every lunch just to unlock a list.
It also means I can see the mistakes and work around them. When Gemma called a recorded video a live interview, I could replay the exact same input as many times as I wanted, try a fix, and check it, at no cost. That is how the "catch needs its words" rule came out of the first trip.
And it runs on what I have. An old Quadro card with 4 GB of memory, no API key, no bill. For someone earning from bounties in Vietnam, that is the difference between a tool I use every day and one I turn off.
My Agent Session
I built this with Claude Code as a pair. It wrote most of the code from my description and tested it on my machine; I chose what to build, walked, and decided which of its ideas sounded like me.
Prize Categories
- Best Use of Gemma



Top comments (0)