What I Built
Trail Bird Journal is an offline bird-song journal that runs entirely on my phone. You tap Record (or upload a clip), an open bird-sound model names the species it hears, and the app writes a short field-journal entry and saves it with the audio. A second page keeps a growing life list of every bird you've identified.
It is for the person who walks past a hundred birds a day and has never known their names, and especially for the places where birdwatching actually happens: trails, woods and wetlands where there is no signal. A cloud app is useless there. This one works with airplane mode on, and your recordings never leave the phone.
The point is to get you outside and keep you standing still. The reward for listening is a name, an entry in your journal, and one more bird on the list.
A real entry from my outdoor test:
Date : 10/10/2026
Place : Near an open ground
Species recognised : rock dove
Demo
Code
Trail Bird Journal
An offline bird-song journal for your phone. Record or upload a clip, and the app identifies the birds with BirdNET and writes a short field-journal entry. No internet, no cloud, no account. Built for the "Touch Grass" open-source AI challenge.
How it works
- Browser page records audio (or you upload a file) and sends it to a local FastAPI server.
-
ffmpegconverts it to 48 kHz mono WAV. - BirdNET (ONNX backend) lists the species it hears with confidence scores.
- The journal sentence is first built by plain Python from those detections, so it can only state what was detected. An optional small local LLM (Gemma 3 1B, via llama.cpp) may rewrite it, and a validator rejects any rewrite that adds facts, so the template text is used instead.
- Each session is saved as a JSON file plus the audio clip. History and Birds Discovered pages…
Everything is in the repo: a FastAPI server, a small pipeline, and four hand-written pages with no frameworks and no CDN. No model weights are in the repo; the README says where to download them.
How I Built It
I built and ran all of it on an Android phone (iQOO Z9, 8 GB RAM) using Termux, a proot Ubuntu, and opencode as my coding agent.
Browser page (phone)
-> FastAPI server (localhost, inside proot Ubuntu)
-> ffmpeg: any audio -> 48 kHz mono WAV
-> BirdNET (ONNX backend): species + confidence per 3-second window
-> journal.py: builds the entry text, optional LLM rewrite, validator
-> saves entries/.json and audio/.wav
-> optional: llama-server (native Termux) running Gemma 3 1B (Q4_K_M)
Bird identification: BirdNET (Cornell Lab of Ornithology and Chemnitz University of Technology), run through its ONNX backend. It scores each 3-second window of the clip against thousands of species.
Journal writer: Gemma 3 1B Instruct (4-bit GGUF) via llama.cpp, about 11 tokens per second on the phone.
No training. The project wires two open models together and builds the safety net around them.
The design rule that came out of the project: the language model never gets to know anything about birds. Plain Python builds a sentence from the detections, and the model may only rephrase it. That rule exists because of what went wrong, which is most of the story.
Challenges, and what I changed
- TensorFlow doesn't exist for my Python. My proot Ubuntu ships Python 3.14. pip install birdnet-analyzer failed because it needs TensorFlow, which had no wheel for 3.14. The newer birdnet package also defaults to a TensorFlow backend, so that failed too. The fix was its ONNX backend (model version 3.0), which needs no TensorFlow at all. Open model formats let me route around a packaging dead end instead of giving up.
- A crash caused by a missing if name == "main". BirdNET starts worker processes, and Python 3.14 re-imports the script in each one, so without the main guard everything died with an unhelpful "child process exited with exit code 1". The guard plus n_workers=1 (to keep RAM low) fixed it.
- Two tools disagreed about the same clip. My first test clip scored Eurasian Skylark at 0.91, 0.95 and 0.90 across its three segments. A web bird-ID app said Daurian Redstart. BirdNET's top ten never included the redstart. I listened to it: a long, continuous, rising song, which matches a skylark, so I trusted the model whose scores I could see. The lesson is that every ID needs a confidence label, so the app now says "clearly identified", "likely" or "possibly, not certain" based on the score.
- My 1B model lied to me. First test, I asked Gemma for "one fact about the bird". It told me skylarks nest "in the eaves of buildings". They nest on the ground in open fields. When I removed the fact request and tightened the prompt, it still invented things I never gave it: a "forest clearing" (the place was "unknown"), "bright yellow plumage", "a red bird", "a beautiful, melodic sound". On a noisy 80-second forest clip it simply repeated my prompt's format back at me. A 1B model copies the shape of whatever you show it and fills gaps with plausible fiction. What worked: Python decides everything factual (species, hedge words, segment counts). Python writes the entry text itself. This is the default. Gemma may rewrite it, but a validator rejects any rewrite that adds banned words (colors, habitats, behaviors), drops a species, or loses a "possibly" hedge. If the rewrite fails validation, or the LLM isn't running, the app uses the template entry and labels it "template" instead of "local LLM". So the bird ID is never the LLM's problem, and the LLM can't make the journal wrong.
- The counting bug: "23 of 3 segments". On an 80-second clip with 27 windows, my summary said a nuthatch was "detected in 23 of 3 segments". The total was hard-coded from my 8-second test clip. I fixed it by counting unique start times, and later hit a cousin of the same bug: the template once printed "heard 3 of Alauda arvensis segments", where the scientific name had landed in the slot for the total. Small bugs, but a journal that states wrong numbers is worse than no journal.
- 8 GB of RAM is not much. llama-server, code-server and opencode together do not fit. Android killed sessions until I started working in phases: build with the editor and agent, then close them, then test with only BirdNET running, then add the LLM last. A 4-bit 1B model needs about 1 GB, BirdNET about the same, and the OS takes the rest.
- The agent wrote a frontend it never ran. I used opencode to generate the backend and pages. The Python mostly worked. The browser side did not, because the agent had only checked syntax and never loaded a page. What went wrong: Pages linked style.css instead of /static/style.css. Served at /capture, the browser asked for /style.css and got a 404. The Play button was a plain link, so the browser downloaded a file named recording.wav.json containing {"detail":"Audio not found"}. The server looked for audio in one folder while the pipeline saved it in another. Buttons did nothing. The forms were hidden with inline styles and the script that should show them never ran, so Skip and Continue just reloaded the page with a stray ?step=1 in the URL. The server crashed at startup calling a helper that no longer existed. Uploads failed with a 422 because the page sent the file under a different field name than the server declared. I fixed these by hand with small patches, then replaced the capture page with a compact hand-written one, since patching on top of dead code wasn't working.
- The microphone, and a format BirdNET can't read. Recording first failed with "Permission denied" because the browser stores mic permission per address and I was on 127.0.0.2 instead of localhost. Once the mic worked, the recording came back as WebM, which BirdNET cannot read. Rather than depend on ffmpeg inside proot, the page now decodes the recording in the browser, resamples it to 48 kHz mono, and uploads a real WAV, the same kind of file as my working uploads. I took it outside When and where: 10/10/2026, 9:43AM, near an open ground Setup: phone in airplane mode, llama-server and the app running in Termux Clips recorded: 4, each 8 sec What it got right: rock dove Airplane mode check: did everything work offline? yes Why Does Open Innovation Matter? It works where closed APIs can't. The places worth birdwatching have no signal. A hosted model would have meant a recording you can't process until you get home. Open weights and local inference mean the answer arrives while the bird is still singing. Private by construction. Field recordings carry your location, your voice, and sometimes strangers' conversations. Nothing leaves the phone, so there is nothing to leak. Swappable without permission. When TensorFlow had no build for my Python, I switched BirdNET to its ONNX backend and carried on. I can swap Gemma for another small model, or change the bird model, by changing one path. Inspectable, so I could constrain it. Because I could see BirdNET's raw scores and run the LLM on my own hardware, I could find exactly where the 1B model made things up and build a validator around it. With a black-box API I could only have hoped the prompt held. Free to run. No account, no key, no per-request cost, and no bill for walking around a pond for an hour. Where open beat closed here was not raw quality (a large hosted model would write nicer sentences). It was availability, privacy, and the freedom to fix things myself. Honest caveats: BirdNET's model weights are licensed for non-commercial use, so this project is a personal and educational one. The birdnet Python package itself is MIT-licensed. Gemma is open-weight under Google's terms of use, not an OSI open-source license. Prize Categories Goggle gemma
Top comments (1)
tr.ee/dev-to