This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
๐ฑ What if an app's best outcome was that you stopped using it?
That is the idea behind Outbound: a small, gamified outdoor quest app built around local, open-weight AI. Not an infinite feed. Not another chatbot to keep talking to. Just a little help choosing something worth doing in the real world.

Figure 1. A starting point, not a destination: the actual Outbound interface with automated demo state.
You start with interests, goals, things you dislike, and a few practical reminders. Before an outing, you tell Outbound how much time you have, how you feel, your energy level, and optionally your starting point.
๐ฏ It gives you three possibilities. You choose one. Then you can put the phone away.
The activities range from gentle movement and noticing nature to tiny creative exercises, low-pressure social activities, and useful errands. There are 120 distinct catalog activities, not 120 rewrites of โgo for a walk.โ

Figure 2. Choose โ go do it โ reflect. These are automated UI fixtures, not photographs or evidence of a completed outdoor trip.
On return, you can mark the quest complete, partial, or skipped; optionally say whether it felt worthwhile; and leave a short note. โBring water next timeโ can become a future preparation reminder. Missing ratings stay missing rather than becoming dislikes.
๐ The rewards acknowledge participation. Badges include First Step, Three Real Days, Curiosity Engine, Screenbreaker, and Partial Counts. Local sounds celebrate progress. Optional ElevenLabs read-aloud and reward speech add a friendly voice, while typed notes and browser speech remain available.
This is for the person who thinks, โI should get outside,โ but gets stuck choosing what to do. Someone with fifteen minutes and low energy deserves a useful option too.
๐ฑ The design goal is not more time in Outbound. It is less friction before doing something outside.
Demo
๐ Try Outbound
To explore the full decision trail without affecting your own progress, open Recommendation lab, seed the separate demo workspace, and run a recommendation. The lab distinguishes synthetic examples from personal outcomes.

Figure 3. Named destinations include travel, activity time, and a reserve. โNeighbourhood Greenโ and the displayed source readings are controlled visual-test fixtures, not live place-discovery evidence.
๐บ๏ธ A destination is earned by the data, not invented by the language model.
For an illustrative 30-minute outing, a six-minute walking round trip, twelve-minute sketch, and two-minute reserve total twenty minutes. That fits. A garden with a 28-minute round trip does not fit a fifteen-minute outing, however attractive the model could make it sound.
No precise starting point, no usable route, insufficient time, or after-dark exclusions? Outbound can offer suitable location-independent catalog activities instead. It does not attach a made-up garden or guess a walking time.
โณ This is a CPU-hosted prototype. Cold models and busy public map services can make generation slow. Opening hours, facilities, and access remain unverified; check conditions yourself before setting out.
Code
๐ป Source, setup, and architecture on GitHub
git clone https://github.com/Sahil-Jaiswal-189/Outbound.git
cd Outbound
npm install
cp .env.example .env
npm start
This starts the local web app. The README includes the additional Ollama/TabPFN setup and a single-container option that runs Node, both model services, and SQLite together. Credentials belong in .env or the hosting dashboard, not in Git.
How I Built It
๐งฉ Two models, a transparent selection policy, and clear responsibilities.

Figure 4. Self-hosted inference and storage are separate from external weather, map, route, and optional voice services.
๐๏ธ SQLite remembers; it does not predict
SQLite stores the editable profile, context snapshots, saved recommendations, attempts, feedback, source cache, and structured events. Feedback refers to the saved quest rather than accepting replacement facts from the browser.
Unchosen activities are not failures. Started-but-unfinished activities do not silently become training labels. Partial attempts keep their own status, even though the current completion classifier specifically predicts full completion.
๐ฆ๏ธ Specialists provide facts
The Node backend calls structured sources directly:
- Open-Meteo: outing weather and modeled air quality.
- OpenStreetMap / Overpass: named nearby places.
- openrouteservice: pedestrian round-trip estimates and area names.
Leaflet supplies the map interface. Overpass discovers places; ORS checks routes. Qwen does neither. Source calls are deterministic adapters, not an autonomous tool-calling agent.
A catalog/place matcher builds appropriate activities at supported destinations. Filters then consider time, effort, preferences, explicit supported constraints, weather, and daylight. A nearby shop alone is not a reason to invent a shopping need.
๐ TabPFN estimates outcomes, not destinations
The local Python service uses TabPFN v2 to estimate two separate probabilities: full completion and enjoyment. Its thirteen pre-outing features include available time, mood, energy, goal, locality type, weather, activity type, duration, travel, and physical/social effort. Exact coordinates and raw notes are not predictor features.
Cold start uses a smoothed category baseline. Each target needs enough varied labeled data, including at least thirty earlier training examples and eight later holdout examples. TabPFN is promoted for that target only when its holdout Brier score is lower than the baseline's. One target can use TabPFN while the other remains on the baseline.
๐ฌ This is a preliminary quality gate, not proof that the app improves fitness or that small-sample probabilities are perfectly calibrated.
๐ฒ Selection balances relevance with variety
The recommendation engine combines outcome estimates with goal alignment, bounded mood/place bonuses, and repetition penalties. It creates a diverse slate, avoids recently offered activities when alternatives remain, and uses limited epsilon-greedy exploration in the third slot.
That policy is inspectable code, not a separately trained reinforcement-learning model. It records its candidate pools and conditional selection probabilities. The probabilities describe the app's selection, not a causal estimate of what an activity will do for someone.
๐ฌ Qwen supplies warmth after the decision
Qwen2.5:3b runs through Ollama. It can customize friendly generic titles and reflection prompts after the engine selects activities. Named destination titles, factual steps, timings, and preparation are preserved. Final recommendation reasons are built from saved evidence rather than trusting an LLM-written explanation.
New feedback becomes context for future predictions and relevant reminders. We do not fine-tune Qwen or TabPFN weights after every quest.

Figure 5. The lab makes the decision process inspectable. This screenshot uses a controlled visual fixture; actual model execution is checked separately.
๐ ๏ธ Built to Be Trusted
A quest should come with evidence, not just reassuring language. Named destinations need a mapped place, a successful walking estimate, and enough time for the activity and return. When a provider or model is unavailable, Outbound reports it and uses the appropriate cached-data, baseline, or template fallback. Missing information never becomes an invented destination.
Your progress should survive the app restarting. The single-container deployment preserves SQLite history and model caches on its persistent disk. Supervisor restarts crashed workers, while the model services stay internal and only the web API is exposed publicly.
โ Checked beyond mock responses: 69 automated tests cover the backend, Python predictor, and desktop/mobile workflows. A separate container smoke test runs real Qwen generation and TabPFN evaluation, kills a worker to check recovery, and verifies that saved history survives a fresh-container redeploy.
Why Does Open Innovation Matter?
๐ The important freedom is being able to change what the system optimizes for.
I can inspect the feature list, change the scoring weights, swap the predictor or local language model, and examine why a quest was selected. Outbound does not depend on asking a hosted chatbot to make an opaque recommendation and hoping the explanation is true.
๐ Local inference keeps the core under my control. In laptop mode, profile/history storage and model execution stay on the laptop. In the container deployment, they stay on the server I operate rather than being sent to a hosted inference API. Once weights are downloaded, the model computations themselves can run without an internet connection; fresh weather, places, routes, and map tiles still need one.
๐งช Open components let me use different tools for different jobs. A small-data tabular foundation model estimates outcomes. An explicit policy selects a feasible slate. A local language model supplies wording. Neither model has to pretend to be a database, a map, or a safety authority.
๐ธ There is no required hosted-model fee per recommendation. That does not mean the whole app is free to operate: hardware, electricity, hosting, and optional ElevenLabs can cost money.
An important distinction: open-weight does not mean unrestricted licensing. This Qwen checkpoint uses the Qwen Research license; the TabPFN v2 checkpoint uses the Prior Labs License with attribution requirements. Ollama is an open-source runtime. ElevenLabs is a separate proprietary, optional service. Coordinates still go to location providers, and enabled voice/transcription sends its payload to ElevenLabs.
That is where this approach fits better than a closed inference-only integration for my project: control of the decision rules, deployment, and inspectable fallback behavior, not a claim that these models outperform every closed model.
My Agent Session
๐ค I built Outbound with an AI coding assistant, iterating on the catalog, local prediction service, source grounding, audio interactions, deployment, and tests. I do not have a shareable DevRelay recording, so there is no session embed here. The source and architecture document show the implementation; they are not a substitute for a recorded session.
Prize Categories
- ๐ Best Use of TabPFN: local completion/enjoyment prediction, per-target chronological validation, and explicit baseline/hybrid/model modes.
- ๐๏ธ Best Use of ElevenLabs: optional quest read-aloud, friendly reward speech, and reflection transcription.
- ๐ Best Use of Render: the supplied public demo is hosted on Render; the repository also packages the complete self-hosted runtime in one Docker service.
- ๐ป Best Use of GitHub Copilot: used GitHub Copilot during the development of Outbound.
๐ฟ The ambition is small on purpose: make one real-world action easier to choose today.
No invented field-test story. No claim that badges solve habits. The next meaningful evaluation is taking the app outside and learning from genuine outcomes.
Visual credits: the app's park-path photograph is by Tina Devidze on Unsplash, under the Unsplash License. Screenshots show the actual app using automated fixtures; architecture and workflow figures were assembled for this write-up.

Top comments (1)
wow, nice tech
very usefull