A JARVIS-style voice companion for your desktop. Talk naturally. Delegate real work. Control it with your hands. Let it float over everything you do.
License: MIT Platform Built with Electron PRs Welcome Follow on X
Gemini Live is the voice. Hermes Agent is the muscle. Iris is the bridge between you and both.
Demo · Features · Glass HUD · Quick Start · How It Works
🔊 Sound on — Iris is a voice tool. A real session: waking Iris by voice, delegating a task to Hermes mid-conversation, live tool steps streaming in, opening the result hands-free with a finger point, and Glass HUD mode over the desktop.
iris-hud-720p.mp4
Iris is a voice-first desktop companion built on a two-brain architecture:
- 🎙️ Gemini Live handles realtime conversation — you speak, it answers instantly with natural voice, handles interruptions (talk over it, it stops), and uses built-in Google Search for quick facts.
- 🛠️ Hermes Agent does the real work — research, coding, files, terminal, browser automation, email, Notion, anything tool-heavy. Iris hands tasks to Hermes in the background and keeps chatting with you while Hermes grinds.
- 🖐️ MediaPipe hand tracking lets you drive the whole UI in the air — point-and-hold to click anything, open palm to scroll, fist to close. Fully on-device.
- 🪟 Glass HUD mode turns Iris into a transparent, click-through overlay floating above your entire screen — keep coding, browsing, writing, while Iris hovers like a heads-up display.
When Hermes finishes a background task, Iris proactively speaks up: "Quick update — Hermes is back with the result." That's the loop: you talk, Iris routes, Hermes works, Iris reports back.
- Realtime voice via Gemini Live (16 kHz in / 24 kHz out, WebRTC echo cancellation — works fine on laptop speakers)
- "Hey Iris" wake word — local, on-device ONNX inference; nothing leaves your machine while asleep
- Barge-in — interrupt Iris mid-sentence and it yields immediately
- Speech-aware standby — local Silero VAD protects long utterances from the idle timer without using room-volume thresholds; Gemini retains its own automatic VAD for conversation turns
- Voice-driven UI — "open the latest result", "show the steps", "open the failed one", "close it" — fuzzy-matched against what's on screen
- Shared personal context — a bounded snapshot of Hermes
USER.mdandMEMORY.mdkeeps ordinary conversation coherent; brain notes and deeper memory details remain available through sourced retrieval tools
- Background delegation — tasks return a
run_idinstantly; conversation never blocks - Live activity feed — every tool call Hermes makes (browser, code, files, search) streams into the task card in real time with durations
- Full Hermes interaction bridge — native clarification choices/free text, dangerous-command approvals, sudo passwords, and secret prompts surface in Iris. Sensitive values stay UI-only and never enter Gemini.
- Visible voice answers — spoken clarification text fills the on-screen field/choice briefly before Iris submits it to Hermes
- Proactive completion announcements — Iris tells you the moment work finishes and summarizes it out loud
- Intent-aware confirmation gate — Iris reads back an immutable proposal and interprets your next response conversationally, with no required approval phrase. Code still enforces a separate user turn, proposal ID, exact brief, selected Hermes thread, and TTL before dispatch.
- Anti-hallucination guardrails — status and results can only come from real Hermes API responses; "still working" is the only honest answer until then, and tool responses say so explicitly
- One pinned Hermes chat thread — like picking a chat in the Hermes app; no stray sessions
- Session switcher on the main page — chip at the top of the Work Stream lists your Iris sessions; + starts a new thread (named by Hermes, titled by your first prompt, like every chat tool)
- Session-scoped Work Stream — only the selected session’s cards are shown; work in other sessions continues without leaking cards into the current thread
- History restore — close Iris, reopen it, and your past completed runs are rebuilt from Hermes's own session transcript. Nothing is lost between launches.
- Point-and-hold to click anything (~0.3 s) — cards, buttons, toggles, chips; the reticle visibly "charges" while arming
- Open palm — joystick scrolling over whatever region your hand hovers: work stream, comms, step timelines, the reader
- Two open palms — resize the reader
- Closed fist — close the reader
- All tracking is on-device (MediaPipe GestureRecognizer); frames never leave your machine
- Deep-space design system — electric cyan + violet on flat deep navy (deliberately not near-black — see below), glass panels, Inter / Space Grotesk / JetBrains Mono
- The orb — a canvas-drawn arc reactor that breathes with the actual audio level, changes palette per state (listening / speaking / working), and flashes when work is handed off
- Orb micro-expressions — a double-pulse on wake, a soft ripple when your words are locked in, and an orbiting "thinking" swirl in the gap before Iris speaks
- Two voice signatures — your voice renders as sharp radial bars around the orb; Iris's voice as a smooth breathing wave. You can see who's talking.
- Sound design — five subtle synthesized cues (wake, sleep, task sent, task done, approval needed). No audio files; pure tuned Web Audio tones. Toggle in Settings.
- Comet handoff — a particle streaks from the orb to the task card when Gemini delegates, and back when Hermes returns
- Cinematic boot sequence, animated transitions, and a custom app icon rendered from the orb itself
If you have ever built a dark interface and seen faint contour lines — concentric rings or horizontal bands across large dark areas — this section is why Iris does not have them. It is the most transferable thing in this design system, so it is written up in full.
The cause. A display gives you 256 steps per channel, and sRGB spends very few of them near black. A soft gradient stretched across 900px has only a handful of values available to cross that distance, so it cannot render as a gradient at all: it becomes a few wide flat plateaus with a hard 1-level edge between each. Those edges are the contour lines. It gets worse the darker you go, because below the sRGB toe brightness is proportional to code value — one code step at value 7 is a ~14% jump in brightness; at value 21 the same step is ~5%, which is under the threshold where the eye reads it as a line.
This is not a bug anyone has fixed. You can see it in Cursor, in native macOS windows, and in dark video. Polished apps don't solve it — they avoid causing it. Look closely at Linear, Raycast, Xcode or Things: almost no large soft gradients anywhere. Flat surfaces, hairlines, small shadows, and gradients only where they are small and steep.
The rules Iris follows.
- If a gradient crosses more pixels than it has code levels to cross them with,
it is not a gradient — it is a stack of plateaus with visible edges. Make it a
flat fill. The deck background was two window-sized radial washes plus a
--bg-1 → --bg-0ramp; that ramp spanned 2 levels of green over 860px. Invisible as shading, glaring as edges. - Keep the whole palette out of the crush zone.
--bg-0is#0b111c, not the#030509it started as. Everything painted later — vignettes, scrims, overlays — has to respect the same floor, or it drags its own region back down. - Gradients on dark surfaces must be small and steep. A 40px button fade has plenty of levels for its range. A 900px wash does not.
- Depth comes from a step of fill tone plus spacing, not from washes or outlines. Surfaces here nest by tone: background → panel → card. None of them carry a border, because a surface a step lighter than its parent already reads as separate, and saying it twice means the lines are what you notice.
- One light source, and let it be real. The reactor canvas is the only thing that glows. Beware the trap: once the background is flat, a single soft halo becomes the only brightness variation, so a bright centre falling off to dark corners is exactly what vignetting looks like.
What does not work, in case you are tempted:
- A noise/dither overlay. The standard advice, and it is a band-aid. It cannot
remove a staircase that is already baked into a rasterized layer — it only
textures over it, so the low-frequency steps the eye integrates survive. Worse,
backdrop-filteraverages its backdrop, so any dither underneath a blurred glass panel is smoothed away and the banding comes back inside the glass. - Blend modes.
mix-blend-mode: overlayresolves to2 × backdrop × sourcebelow mid-grey, so on a near-black surface its effect is a fraction of one code level. Invisible. - More gradient stops. The steps are a quantization limit, not a stop-count problem.
If you need to measure it. Run-length checks will lie to you: dither makes runs look short while the banding is untouched. The eye integrates along a long straight edge, so average each row across a wide strip of background first (which cancels per-pixel noise), then look for jumps in that row-average. Max local slope is the number that tracks visibility — a smooth 30-level ramp over 600px climbs at ~0.05 levels/px, and anything spiking well above that is a line you can see.
These rules are enforced for future work, not just documented: they live as a Cursor rule in
.cursor/rules/dark-theme-no-banding.mdc, scoped tosrc/styles/**/*.css, so anyone (or any agent) editing the styles gets them automatically. Copy that file into another project to carry the approach over.
The signature feature. Press ⌥H anywhere (or say "enter HUD mode"):
- Iris becomes a transparent, always-on-top overlay across your whole screen
- Click-through everywhere except Iris's own elements — your mouse works normally in the apps underneath
- Collapsible Tasks and Comms chips, camera tile, caption pill by the orb
- Open results in a glass reader without blacking out your desktop
- Gestures get even better here: your hand reticle floats over your entire screen
- Menu-bar tray icon keeps Iris alive in the background (wake/sleep/HUD from the tray)
- Node.js 20+ and npm
- Hermes Agent installed with the gateway running
- A free Gemini API key from Google AI Studio
- macOS, Windows, or Linux (Glass HUD is built and tuned on macOS)
git clone https://github.com/ASHR12/iris.git
cd iris
npm install
npm run devecho 'API_SERVER_ENABLED=true' >> ~/.hermes/.env
echo "API_SERVER_KEY=$(openssl rand -hex 32)" >> ~/.hermes/.env
hermes gateway restartHermes requires a strong key (16+ chars) and refuses to start its API server with a short or placeholder one — this endpoint dispatches terminal-capable agent work. Use the same key in Iris's
~/.iris/.env.
Iris opens an onboarding wizard on first launch — paste your Gemini key (with a live Test button), point it at Hermes (Test that too), pick a voice (with preview), grant mic permission. Everything saves to ~/.iris/.env. No manual file editing needed.
Press ⌥W (or say "Hey Iris" if you enabled the wake word). Ask it something. Then ask it to do something — "check my unread emails and summarize what needs attention" — confirm the brief, and watch Hermes go to work while you keep talking.
Try it with zero keys: toggle demo mode in Settings → Advanced.
Dloads a full fake workspace,Gplays a simulated handoff. Great for screenshots and poking at the UI.
Production build & packaging
npm start # build + launch production bundle
npm run package:mac # macOS .app; uses IRIS_MAC_SIGNING_IDENTITY or "Local Development" when available
npm run install:mac # build + install to /Applications + launch (macOS)
npm run dist:win # Windows distributableLocal macOS packaging post-signs and verifies the bundle with the configured identity. If that certificate is unavailable, packaging falls back to an unsigned local build.
The packaged app reads config from ~/.iris/.env (%USERPROFILE%\.iris\.env on Windows), written by the wizard.
| Input | Action |
|---|---|
| ⌥W / ⌥S | Wake / sleep (global — works from any app, even with the HUD over your work) |
| ⌥H | Toggle Glass HUD (global — works from any app; configurable via IRIS_HUD_HOTKEY) |
| "Hey Iris" | Wake by voice (opt-in, on-device) |
| Top-right buttons | HUD toggle · Settings · hand-tracking toggle · link status |
| Menu-bar icon | Wake/sleep, HUD, show deck, quit |
| D / G | Load demo data / simulate a handoff (demo mode only) |
| Gesture | Action |
|---|---|
| ☝️ Point (index up) | Move the cursor — hold ~0.3 s over anything clickable to click it |
| ✋ Open palm | Hold-to-scroll whatever region your hand is over (high = up, low = down) |
| 🙌 Two open palms | Resize the open reader |
| ✊ Closed fist | Close the reader |
- "What's the latest on the OpenAI news?" → answered directly with Google Search
- "Check my unread emails and tell me if anything needs attention" → read-back → your natural confirmation → Hermes runs it in the background
- "How's that task going?" → real status from the Hermes API, never invented
- "Open the latest result" / "show the steps" / "open the failed one" → UI obeys
flowchart LR
You(("🎙️ You")) <-->|"realtime voice"| Iris["IRIS<br/>(Electron + React)"]
Iris <-->|"16k/24k PCM audio,<br/>tools"| Gemini["Gemini Live"]
Gemini -->|"quick facts"| Search["Google Search"]
Gemini -->|"propose → you confirm →<br/>submit_hermes_task"| Hermes["Hermes Agent<br/>(TUI Gateway)"]
Hermes -->|"JSON-RPC events:<br/>tools, clarify, approval, secure input"| Iris
Hermes -->|"completion event"| Gemini
Gemini -->|"'Hermes is back —<br/>here's the result'"| You
- Electron main owns the Gemini Live session — mic audio streams up, voice streams back, transcripts and state events flow to the UI.
- Gemini routes: quick public facts use Google Search, personal recall uses sourced
search_memory/read_memory_note, UI commands stay local, and real work goes through the immutable two-step dispatch gate. - Hermes runs in the background through its authenticated TUI Gateway WebSocket. Tasks in one chat queue safely; different chats may work concurrently. Every tool step and native interaction request streams into Iris.
- On completion, Iris injects a system event into Gemini so it proactively announces and summarizes the result — then you can open it by voice, mouse, or finger.
- Sessions mirror Hermes: all work lives in one pinned Hermes chat thread; the Work Stream rebuilds from that thread's transcript on every launch.
Pinned models, SDKs & known footguns (for contributors)
| Purpose | Identifier | Where |
|---|---|---|
| Gemini Live model | models/gemini-3.1-flash-live-preview |
electron/main.mjs (GEMINI_LIVE_MODEL) |
| Gemini SDK | @google/genai |
package.json |
| Gesture runtime | @mediapipe/tasks-vision (WASM pinned to same version in useHandControl.ts) |
package.json |
| Wake word | openWakeWord-style ONNX via onnxruntime-web |
public/wakeword/ |
- Live models are a distinct family — a normal
gemini-*chat model will not open a Live session. Keep themodels/prefix. - Gemini Live audio is fixed-format: send 16 kHz PCM, receive 24 kHz PCM.
- Live function calls are synchronous — never block a tool call on Hermes work; return the
run_idimmediately. - Keep the MediaPipe WASM CDN version identical to the installed npm package version.
- MediaPipe WASM + model are fetched from CDN on first gesture use — first run needs network.
Everything is configurable through Settings (gear icon) — the file below is written for you. Power users can edit ~/.iris/.env directly:
GEMINI_API_KEY=... # required — aistudio.google.com/apikey
IRIS_USER_NAME=Ashutosh # what Iris calls you
GEMINI_LIVE_MODEL=models/gemini-3.1-flash-live-preview
GEMINI_LIVE_VOICE=Zephyr # pick + preview in Settings
HERMES_API_URL=http://127.0.0.1:8642
API_SERVER_KEY=<openssl rand -hex 32> # 16+ chars, must match Hermes's ~/.hermes/.env
HERMES_HOME=~/.hermes # optional, auto-detected
IRIS_HERMES_TRANSPORT=interactive # full TUI Gateway protocol; runs_api is legacy fallback
# IRIS_HERMES_CWD=/safe/workspace # optional; unset uses Hermes terminal.cwd (recommended)
# IRIS_HERMES_PROTECTED_PATHS=~/Documents,... # optional protected-path override; Downloads allowed by default
IRIS_HERMES_SESSION=iris-voice # pinned Hermes chat (or use the UI switcher)
IRIS_HERMES_MEMORY_KEY=iris:desktop:your-name # stable long-term memory scope (auto-derived if omitted)
IRIS_WAKE_WORD=true # "Hey Iris" on-device wake word
IRIS_WAKE_SENSITIVITY=balanced # relaxed 20% / balanced 30% / strict 40%
IRIS_SHOW_WAKE_DIAGNOSTICS=false # optional 6-second wake reason/confidence overlay
IRIS_HUD_HOTKEY=Alt+H # global Glass HUD hotkey
IRIS_AUTO_SLEEP_SECONDS=30 # standby after N s of silence (0 = never) — saves ~80% of Live API cost
IRIS_AUTO_WAKE_ON_HERMES=true # Hermes results wake Iris from standby to announce themselves
IRIS_SOUNDS=true # subtle interface sound cues
IRIS_LOAD_TEST_DATA=false # demo modeConfig resolution order: repo .env (dev) → ~/.iris/.env (wizard/packaged) → bundled .env.
electron/ main process — Gemini Live session, Hermes bridge, dispatch
gate, Glass HUD window control, tray, config
liveSessionState / liveToolCoordinator — turn lifecycle,
resume handles, serialized tool execution
hermes* — HTTP/gateway clients, event stream, interactive
transport, dispatch gate, result service
brainIndex.mjs — embedding + BM25 retrieval stack (also a CLI)
runRegistry, memoryService, configStore, rendererBridge,
routingPolicy, approvalPolicy, sleepIntent, windowSecurity
src/
components/ TopBar, CommsPanel, WorkStream, WorkCard, CenterStage,
HudShell, BrainGraph, ReaderOverlay, SessionSwitcher,
SetupPanel, DevicePicker, HermesInteractionPrompt,
ApprovalPrompt, CameraDock, …
hooks/ useAudioPipeline, useHandControl, useWakeWord, useHandoffFx
lib/ audio/PCM helpers, task utils + fuzzy matching, sounds, fixtures
styles/ deep-space design system (tokens → base → deck → overlays → fx
→ hud → brain)
test/ 17 Node test files — gate, Live lifecycle, Hermes contract,
registry, memory, routing, security
scripts/ dev launcher, icon renderer, macOS package/install/sign,
demo-vault generator, live-API harnesses, soak runner
public/wakeword/ on-device "Hey Iris" ONNX models
public/audio/ PCM capture AudioWorklet
build/ app icon (SVG source + renderer) and tray assets
demo-obsidian-vault/ 164 generated fictional notes for demoing the Neural Map
sidecar/ legacy Python Live reference (off unless explicitly enabled)
| Layer | Tech |
|---|---|
| Shell | Electron (transparent frameless window, global shortcuts, tray) |
| UI | React 19 + TypeScript + Vite, custom CSS design system |
| Voice | Gemini Live API (@google/genai), Web Audio (capture, playback, metering) |
| Agent | Hermes TUI Gateway JSON-RPC/WebSocket + local session API |
| Gestures | MediaPipe Tasks Vision GestureRecognizer (GPU, on-device) |
| Wake word | onnxruntime-web, openWakeWord-style mel → embedding → classifier pipeline |
PRs and issues are very welcome. Good first areas: new orb themes, additional gesture mappings, Windows/Linux HUD polish, sound design, more voice UI commands.
npm run dev # hot-reload dev loop
npm run build # typecheck + bundle
npm test # gate, Hermes contract, memory, routing, persistence and security
npm run verify # tests + Electron checks + build + Python reference compile
npm run test:live-search # real Gemini Search must survive aggressive standby
npm run test:live-sleep # explicit sleep must resume cleanly into a new turn
npm run test:live-speech # active speech must block standby until speech ends
npm run test:hermes-wake # measure Hermes completion wake and announcement
npm run soak # 8-hour Electron RSS/heap and wake/HUD lifecycle samplerPlease keep the two golden rules: voice must never regress because of gestures, and Iris never invents Hermes results — status comes from the API or it doesn't exist.
Release history lives in CHANGELOG.md.
Found a vulnerability? Please report it privately — see SECURITY.md. Iris dispatches terminal-capable work to Hermes, so anything that can influence a dispatch is treated as a security issue rather than a UX one.
- Your Gemini key and Hermes key live in mode-
0600~/.iris/.env, are never returned to renderer state, and are never committed. - Camera frames and wake-word audio are processed entirely on-device and never uploaded.
- Conversation audio goes to Gemini Live while Iris is awake. Wake-word audio stays local while asleep; a brief silent connection may renew the Gemini resumption handle during long standby.
- A bounded snapshot of Hermes
USER.mdandMEMORY.mdis sent to Gemini at session setup. Brain-note content is retrieved only when relevant; semantic indexing/querying also uses Gemini embeddings when enabled. API_SERVER_KEYmust be a strong secret (Hermes enforces 16+ chars and refuses weak keys — the endpoint dispatches terminal-capable agent work). Generate one withopenssl rand -hex 32and use the same value on both sides.
MIT — build wild things with it.
Built by Ashutosh Shrivastava — AI consultant & tech storyteller, 1M+ weekly reach on X
🌐 Website · 𝕏 / Twitter · YouTube · LinkedIn · Instagram · Threads · GitHub
Crafted with Gemini Live, Hermes Agent, and an unreasonable love for arc reactors.
⭐ If Iris made you grin, star the repo — it genuinely helps.