Stars
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python depend…
[ACL 2026] Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
Real-time, low-latency AEC at 48 kHz — designed for efficiency and low CPU usage.
Lightweight on-device keyword spotting engine for iOS using CoreML and real-time audio streaming.
Elucidated Text-To-Audio (ETTA) is a SOTA text-to-audio model with a holistic understanding of the design space and trained with synthetic captions.
A LiveKit TTS plugin for self hosted open source models. Expose any model through a WebSocket compatible API (low latency) and swap engines without changing agent code.
"ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"
The retrieval layer for production AI systems. Lightning-fast (<10ms) search without vector databases. Built for browser, edge, on-device, and cloud.
A standalone desktop/smartTV overlay that translates system audio into 3D Sign Language animation in real-time.
Open-source American english TTS model. 6 voices and a high performance inference library for Apple Silicon.
[SIGGRAPH 2026] Pixal3D: Pixel-Aligned 3D Generation from Images
Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).
[ECCV 2026 Oral] Implementation of "Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length"
Self hosted, real-time digital human agent platform. Build voice-first AI agents with WebRTC, persona memory, tools, RAG, and optional digital-human video.
FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST.…
Browser-based text-to-speech powered by OmniVoice. Runs entirely locally via WebGPU and WebAssembly.
Open source video conferencing app powered by LiveKit. Built with Django and React.
Self-hosted DTLN noise suppression plugin for LiveKit Agents — no cloud API, no per-minute fees
Building actual open source including dataset Multilingual TTS more than 150 languages with Voice Cloning.
A framework for efficient model inference with omni-modality models
🎙️ VoxSherpa TTS Offline Neural Text-to-Speech Engine for Android ⚡ Sherpa-ONNX powered 🔊 Natural voice synthesis 📱 Fully offline processing 🚀 No cloud • No limits
Vietnamese TTS with instant voice cloning • On-device • Real-time CPU inference • 24kHz audio quality • Chuyển văn bản thành giọng nói tiếng Việt • Text to speech tiếng Việt • TTS tiếng Việt