Stars
An autonomous novel writing pipeline, by Hermes Agent
⚒ Evolutionary self-improvement for Hermes Agent — optimize skills, prompts, and code using DSPy + GEPA
Native web workspace for Hermes Agent — chat, terminal, memory, skills, inspector.
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
An Open Phone Agent Model & Framework. Unlocking the AI Phone for Everyone
Local-first chat history analyzer with AI. | 本地优先的 AI 聊天记录分析工具
Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.
An Open-Source Project to Unify Audio Processing and Generation
High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static bin…
π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.
A collection of DESIGN.md files analysis by popular brand design systems. Drop one into your project and let coding agents generate a matching UI.
一个基于小智、xiaozhi-server的Android、IOS语音对话应用,支持实时语音交互和文字对话。现在是flutter版本,打通IOS、Android端。请同志们动动小手,点点小星星,予以鼓励。
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages. No accounts, no API …
Melonking906 / sequelpro
Forked from sequelpro/sequelproSequel Pro with Sequel Ace bits backported into it.
Unlimited-length talking video generation that supports image-to-video and video-to-video generation
An AI skill that provides design intelligence for building professional UI/UX across multiple platforms.
Create and share 3D architectural projects.
Official PyTorch+CUDA Full-functional Web Demo for MiniCPM-o 4.5
VideoLLM-online: Online Video Large Language Model for Streaming Video (CVPR 2024)
[ICLR 2024] Efficient Streaming Language Models with Attention Sinks
🤯 LobeHub is your Chief Agent Operator, organizing your agents into 7×24 operations by hiring, scheduling, and reporting on your entire AI team.
FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI…
OpenUI let's you describe UI using your imagination, then see it rendered live.
视频转字幕、字幕翻译、AI 配音与声音克隆、字幕烧录——免费开源的一站式桌面工具。基于 Whisper / FunASR 等本地模型离线语音转文字,批量处理 + 全平台 GPU 加速,跨 Windows / macOS / Linux。Free, open-source desktop app to generate, translate, dub & burn video subtitle…
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs