-
Fudan University&Sun Yat-Sen University
- Shanghai, China
- chengqy2019@foxmail.com
- @cheng_qinyuan
- https://xiami2019.github.io/
Lists (13)
Sort Name ascending (A-Z)
Stars
A method for multi-reference image-grounded video captioning.
Official implementation of the paper "Acoustic Music Understanding Model with Large-Scale Self-supervised Training".
Gym-Anything: Turn any Software into an Agent Environment
Reconstructing Big Tech T2I/T2V captioning pipelines for production and research
[Official Repo] JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
🎨 The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, …
Elemental Diagnosis of Generalist Mobile Manipulation Policies
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
VidaForge: Building a Video Foundation Model Pretraining Data Pipeline from Scratch in an Academic Lab
Production-ready MoE load balancing via real-time expert replication
MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
This package contains the original 2012 AlexNet code.
Official PyTorch+CUDA Full-functional Web Demo for MiniCPM-o 4.5
the optimized, token-efficient version of the leaked Claude Fable 5 / Mythos 5 system prompt. Re-engineered into clean Markdown for universal execution on Gemini 3.1 Pro, ChatGPT 5.6, and advanced …
LAION MOSS Local-Transformer 1.5 voice-acting model (4.55B, 48 kHz) — inference, fast-inference guide, best-of-64 audio demos
MOSS-TTS-v1.5 8B voice-acting model — inference (single + fast batched), training recipe, evals, and throughput benchmarks. Apache-2.0.
Native End-to-End Full-Duplex Spoken Language Model
Target-level automatic benchmark for raw-input Chinese news TTS pronunciation accuracy
SpaceXAI's coding agent harness and TUI. Fullscreen, mouse interactive, extensible.
Open Source framework for voice agents, multimodal apps, and realtime AI. Maintained by Daily and the community.