DEV Community

Nayananshu Garai
Nayananshu Garai

Posted on

Sky-Whisper

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

SkyWhisper turns stargazing from a looking activity into a listening one. You open the app at home, it computes exactly what's over your sky tonight, writes you a ~90-second narrated guide, and synthesizes it into audio you download. Then you go outside, put the phone face-down, lock the screen β€” and just listen through your earbuds, with lock-screen controls handling the rest.

It's for everyone who has ever held a glowing phone up to the night sky and ruined the very thing they went out to see: your eyes need 20–30 minutes of darkness to fully adapt, and one glance at a bright screen resets that clock to zero. Every other stargazing app is the screen. SkyWhisper's entire thesis is that in the field, the screen's only job is to stay off β€” if you never unlock your phone, you used the product perfectly. That's about as "touch grass" as software gets: the output isn't content to consume, it's instructions to look up.

Demo

Live: https://sky-whisper.onrender.com (free tier β€” first load after idle takes ~30–60s while the container wakes)

The 60-second try-it script:

  1. Allow location (or type coordinates) β†’ Prepare My Sky β†’ watch the ritual stages light up with real server progress, not a fake timer.
  2. Press play on the countdown β†’ flip your phone face-down β†’ control it from the lock screen.
  3. Tap the πŸŽ™ button and say "describe the sky" β€” or "where can I find Vega?" β€” and hear the agent answer from live data.
  4. Open any pack's "Tonight's Narration" card: it tells you exactly which model narrated it β€” or precisely why the template did instead. No black boxes.

Code

https://github.com/N-Garai/Sky-Whisper

How I Built It

The one architectural decision everything hangs on: coordinates are computed, never generated. A deterministic ephemeris stack (Skyfield + JPL data, CelesTrak + SGP4 for satellites) calculates Sun, Moon, 5 planets, 18 named stars, and 12 constellations for your exact spot and second. The language model only ever receives those finished facts β€” and every number it utters is checked against them before audio renders. A closed model guessing "what's above you" would be unfalsifiable and occasionally confidently wrong; open math can be diffed against Stellarium by anyone.

The model chain is built for a world where model IDs die. During this build, Groq retired one default ID and NIM returned 410 Gone on another β€” I watched it happen in the logs. So each rung carries a fallback ID (switched only on 404/410-class errors), each rung gets 55 seconds, the whole walk gets 150, every failure is recorded with its reason, and the deterministic template always answers. The pack JSON β€” and the UI badge on it β€” names the winning rung or explains exactly why the template did. Model rot can't silently break this product.

Mastra (@mastra/core, Apache-2.0) orchestrates narration: a docent agent with real tools (get-sky-snapshot over localhost, get-body-direction for any named body), a facts β†’ narrate β†’ validate workflow, per-rung endpoint pinning, and a zod narration envelope β€” with the Python validator re-gating everything as the final word. Voice questions ride the same chain with conversation memory, answered as speech.

Free-tier engineering: the whole serving footprint peaks around ~320MB of 512MB; the Node harness spawns only when a model rung actually runs and is kill-guarded so hung runs can't orphan processes into an OOM; packs cache on 15-minute sky buckets so retries serve instantly; the kernel and satellite data warm at boot.

The frontend: (React 19, Vite, Tailwind v4, Framer Motion, Three.js living starfield, installable PWA with offline pack cache) also ships two things I'm proud of: a drag-to-turn 360Β° sky map with a live facing readout, and a blackout pointer mode β€” pure-black OLED overlay where compass, vibration, and rising pitch guide your aim onto a star without emitting a photon that matters.

Testing: 136 automated tests, hermetic (no network, no keys) β€” ephemeris math, anti-hallucination rejection (including a test pinned from a real production failure where a model emitted raw altitudes), chain fallthrough, kill-on-timeout, and full HTTP integration.

Why Does Open Innovation Matter?

Four concrete ways, all demonstrated in code rather than claimed in prose:

  1. Open coordinates are a correctness feature. My cross-check table (in docs/accuracy.md) diffs Skyfield against astropy's independent implementation: 0.00Β° across the board except a 0.64Β° lunar-elongation residual I document and explain. Try auditing a closed API's sky positions that way.
  2. Swap-ability is a one-line demo. Hosted Gemma via free key β†’ local Ollama via one variable. A closed-API product cannot offer "run my narrator on your laptop with zero code changes"; the open-weight one does, and location history ("where you stargaze") never has to leave your machine.
  3. It costs nothing to run. Free Render + free model tiers + free TTS tier + free data sources = $0/month, every line auditable. Open is what makes the hobby budget possible.
  4. It survives the real world. When providers retired model IDs mid-build, the open, inspectable stack meant I could see exactly what broke (HTTP 404/410 in plain text) and route around it in the open β€” instead of a black-box outage I couldn't diagnose.

Prize Categories

  • Best Use of Render β€” single free web service serving API + PWA; cold-start honesty, ephemeral-by-design storage, measured 320MB footprint on 512MB.
  • Best Use of Gemma β€” open-weight Gemma narrates from computed facts only (never coordinates), hosted free-tier path plus one-variable local Ollama swap, model recorded per pack.
  • Best Use of ElevenLabs β€” pre-trip TTS rendering with an honest transcript-only mode when unkeyed; short voice replies synthesized on demand.
  • Best Use of Mastra β€” docent agent with snapshot/direction tools, factsβ†’narrateβ†’validate workflow, per-rung orchestration, memory-carrying voice loop.

Top comments (0)