DEV Community

Brage 1025
Brage 1025

Posted on

Who's Singing?: An offline bird-call identifier

Hacktoberfest: Maintainer Spotlight

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built:

Who's Singing? is an offline bird-call identifier for your walks. You record 30~60 seconds of audio on your phone, run one command, and get a list of which birds were singing, how much to trust each one, and a short note written by a local LLM.

The idea is to keep the screen the shortest part of the experience. You spend the time outside listening, walking with your ears open and your phone ready to record in your pocket. The the laptop is only checked at the end, usually when you're done for the day. It's for anyone who has stopped on a trail and wondered "what is that?". It works even at your hut or campsite where there's no signal, because nothing in it needs the internet.

Who's singing:

  Great Tit                  confident  best  66.2%  (4 detections)
  Great Spotted Woodpecker   likely     best  80.1%  (2 detections)
  Eurasian Blue Tit          possible   best  41.2%  (2 detections)

Field note (written by qwen2.5:3b):
  On this walk, I clearly heard a Great Tit, four times. I also probably
  heard a Great Spotted Woodpecker, twice.
Enter fullscreen mode Exit fullscreen mode

Demo:

To try it yourself, you need a bird recording of your own (the repo doesn't include audio):

git clone https://github.com/Brage1025/whos-singing.git
cd whos-singing

python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -r requirements.txt

ollama pull qwen2.5:3b           # For the field note

python whos_singing.py my-walk.mp3 --lat 59.91 --lon 10.75
Enter fullscreen mode Exit fullscreen mode

Swap my-walk.mp3 for a recording from your phone, and the coordinates for where you made it. The location is what lets BirdNET leave out species that don't live near you.

Code:

GitHub logo Brage1025 / whos-singing

My submission for the Hacktoberfest 2026 Open-Source AI Challenge, Week 1: Touch Grass

Who's Singing?

An offline bird-call identifier for your walks. Record 30~60 seconds of audio on your phone, run one command, and get a list of which birds were singing, how much to trust each one, and a short note written by a local LLM.

Built for the Hacktoberfest Open-Source AI Challenge, Week 1: Touch Grass.

Demo: Watch it run with Wi-Fi off

The screen is the shortest part of the experience: you spend your time outside listening, and only check the laptop at the end.

Why open-source AI?

Two open models do the work, both are running on your own machine:

  • Works without internet. After a one-time download of the models, no internet is needed. Tested with Wi-Fi turned off.
  • Your recordings stay with you. Audio, locations, and your bird list are never uploaded anywhere.
  • Free to run. No API keys, no per-request costs and no token limits.
  • Swappable…

How I Built It:

Two open models run on my own machine:

  1. BirdNET-Analyzer (Cornell Lab of Ornithology and Chemnitz University of Technology) does the identification. It scores each 3-second slice of audio against thousands of species. I pass in my latitude, longitude and the week of the year, and its species list shrinks from 6,522 to about 240, so a recording from Oslo stops returning sunbirds and cuckooshrikes from other continents.
  2. A small local LLM through Ollama writes the one-line note. I defaulted to qwen2.5:3b.

Around them I wrote a short Python script that:

  • merges the 3-second detections into one row per species,
  • assigns a confidence tier (confident, likely, possible), because a single high score on one window is weak evidence and repeated detections are stronger,
  • sends only confident and likely birds to the LLM,
  • checks the note in code and replaces it with a plain template if it's wrong,
  • logs each session to a local CSV, so over time it builds a private bird list.

The LLM kept getting it wrong:

This was the most useful part of the project. The identification worked from the start. The note took five rounds:

  1. Given only species names, the model invented facts. It gave a Great Tit "brightly colored heads" and described a European bird's call for a North American one.
  2. After I forbade facts, the note got clumsy and hedged a confident bird with "probably".
  3. After I added an example to the prompt, the model copied it. It reported a Wren, which was never in the recording.
  4. Then it repeated a real bird ("I clearly heard a Great Tit... I probably heard a Great Tit").
  5. So the code now rejects any note that misses a bird, repeats one, names one that wasn't detected, or uses "clearly" or "probably" wording that doesn't match the tiers. Certainty wording is now decided in code to ensure quality and stability, not by the model.

I then compared two models on the same recording with the final prompt and check, five runs each:

Model Passed the check
qwen2.5:3b 5 of 5
llama3.2 (3B) 1 of 5, the other four fell back to the template

That's one recording and one prompt, so it's a small test, not a benchmark. But it took only minutes, and switching the winner was as easy as changing one flag: --model.

How it did on real audio:

I tested on three recordings: an old one of my own and two from a friend who prefers to stay anonymous. Only one of the friend's recordings has ground truth, because she knew what was there. That one was recorded just outside Oslo. She saw the Great Tits and Blue Tits, and heard the woodpecker without seeing it. So all three detections were real.

BirdNET found everything with no false detections. The weak spot was my own tiers. The Blue Tit was a real bird, but two detections with a best score of 41% put it in possible, so the note left it out. I'd rather the tool under-claim than confidently name a bird that wasn't there, so I kept that threshold, but it's a trade-off I made by judgment, and I didn't tune it on a single recording.

I also owe an honest note: I didn't get out on a walk myself this week, so I haven't field-tested it outdoors. Everything above was run on recordings, offline.

Why Does Open Innovation Matter?:

  • It runs with no internet. After a one-time model download I turned Wi-Fi off and ran the full pipeline, including the note. That's what makes a bird identifier usable on a trail or out on field work with no signal.
  • Recordings stay on my machine. Audio, locations and my bird list are never uploaded. So I can decide if it's nobody else's business or not.
  • It costs nothing to run. No API key, no per-request charge and no token limits.
  • I could see and fix the weak link. When the small model invented a Wren, I could inspect the prompt, change the model, add a validator and compare alternatives side by side. With a closed API I'd have received the same wrong note with no easy way to tell why.
  • The two parts are separately swappable. BirdNET handles identification and the LLM handles wording, so I can replace either without touching the other.

One caveat on "open": BirdNET's model is open-weight but carries a non-commercial license, so it's free to inspect and use for a project like this, not free of conditions. The code in my repository is MIT.

Special thanks to my dear friend, who prefers to stay anonymous, for the field recording and for telling me what she actually saw and heard.

Top comments (0)