-
Fudan University&Sun Yat-Sen University
- Shanghai, China
- chengqy2019@foxmail.com
- @cheng_qinyuan
- https://xiami2019.github.io/
Lists (13)
Sort Name ascending (A-Z)
Stars
Official PyTorch+CUDA Full-functional Web Demo for MiniCPM-o 4.5
the optimized, token-efficient version of the leaked Claude Fable 5 / Mythos 5 system prompt. Re-engineered into clean Markdown for universal execution on Gemini 3.1 Pro, ChatGPT 5.6, and advanced …
LAION MOSS Local-Transformer 1.5 voice-acting model (4.55B, 48 kHz) — inference, fast-inference guide, best-of-64 audio demos
MOSS-TTS-v1.5 8B voice-acting model — inference (single + fast batched), training recipe, evals, and throughput benchmarks. Apache-2.0.
Native End-to-End Full-Duplex Spoken Language Model
Target-level automatic benchmark for raw-input Chinese news TTS pronunciation accuracy
SpaceXAI's coding agent harness and TUI. Fullscreen, mouse interactive, extensible.
Open Source framework for voice and multimodal conversational AI
[SIGGRAPH Asia 2026] DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
The open-source CapCut alternative
This is a real-time conversation project powered by a VoxCPM-based streaming TTS model.
Zero-dependency browser video editor that AI agents can drive — JSON timeline, MCP + REST, live-reloading UI
MuScriptor is a multi-instrument music transcription model developed by Kyutai and Mirelo.
Plugin for agents to interact with Chatcut
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
MOSS-Transcribe-Diarize 0.9B is an open-source SOTA end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness.
An end-to-end framework for multi-speaker transcription that jointly models who spoke, when, and what.
A ComfyUI custom node designed for LTX 2.3 MSR (Multiple-Subject-Reference) LoRA workflows.
Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and VPS.
Talk to 峰哥 — 克隆任何人的声音和性格,实时语音对话,工程延迟 < 1 秒 | Clone anyone's voice & personality for real-time conversation. < 1s engineering latency.