Stars
Guide on how to build your own scanning tunneling microscope based on John Alexander's piezo scanner approach.
A TensorFlow based wake word detection training framework using synthetic sample generation suitable for certain microcontrollers.
Skills for controlling Android, iOS and cloud phones
Free offline AI video dubbing studio for Windows — voice cloning, translation, subtitles & on-screen-text localization. 100% local, one native .exe, zero Python.
Open-source realtime voice agent server in Go with WebRTC (WHIP), barge-in, streaming STT/LLM/TTS pipelines, plugin system, multi-language SDKs, SIP telephony, ESP32 support & fully local mode.
Inflect is a lightweight, high‑quality text‑to‑speech model designed to deliver surprisingly natural audio with a minimal footprint. It’s an active work‑in‑progress focused on fast iteration, stron…
Chrome extension to download YouTube video/audio/subtitles by capturing the player stream. No yt-dlp.
Native End-to-End Full-Duplex Spoken Language Model
An open ecosystem of parametric human models and perception stacks, starting with GNM Head.
A curated collection of Claude Code skills, organized by category
Project page for ScenA: Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors
MOSS-Transcribe-Diarize 0.9B is an open-source SOTA end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness.
ResearchStudio: Our AI co-author, from research problem to final publication.
Modern robot motion planning library based on Pinocchio.
Community model zoo for Apple Core AI (iOS/macOS 27): 62 models — LLM, VLM, OCR, ASR, TTS, image/video/music gen, forecasting — each gated against its source model and shipped with the recipe that …
Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.
The spec-driven development system where specs and tasks live in your repo and never drift from the code.
[ICASSP 2026] Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
Whisfusion: Parallel ASR Decoding via a Diffusion Transformer
This project's aim is to calculate localized coordinates of a sound source using three ESP32 microcontrollers and digital microphones.
Awesome-LLM: a curated list of Large Language Model
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
General-purpose agent in one static Go binary. ReAct loop, ACP server for IDEs, OpenAI-compatible REST API with embedded web UI, Telegram gateway, cron scheduler, long-term memory, context compacti…
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
Scripts for agents, shared between my repositories.
AI-powered UGC pixel-art JRPG: input any address or photo, describe a style, and get a complete interactive world with AI-generated local NPCs.
PTY-backed Claude Code wrapper that provides claude -p compatible output