Stars
Text-audio foundation model from Boson AI
Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable…
[ICML 2026] ByteDance's All-in-One Video Generation Model for Human-Object Interaction Video Generation
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine
Code and results accompanying the paper "Refusal in Language Models Is Mediated by a Single Direction".
🏛️ 三省六部制 · OpenClaw Multi-Agent Orchestration System — 9 specialized AI agents with real-time dashboard, model config, and full audit trails
The most powerful local music generation model that outperforms almost all commercial alternatives, supporting Mac, AMD, Intel, and CUDA devices.
[ICLR 2026] ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
[CVPR 2025] Official repository for "From Poses to Identity: Training-Free Person Re-Identification via Feature Centralization"
SoulX-FlashTalk is the first 14B model to achieve sub-second start-up latency (0.87s) while maintaining a real-time throughput of 32 FPS on an 8xH800 node.
GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation.
[AAAI 2025]👔IMAGDressing👔: Interactive Modular Apparel Generation for Virtual Dressing. It enables customizable human image generation with flexible garment, pose, and scene control, ensuring high …
A curated list of awesome research papers, projects, code, dataset, workshops etc. related to virtual try-on.
📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程
AI for All: The First Systematic Vibe Coding Tutorial | From Zero to Full-Stack, Bring Your Ideas to Life | Live at: www.vibevibe.cn ;全民AI学习第一课,首个系统化 Vibe Coding 开源教程 | 零基础到全栈实战,让人人都能借助 AI 实现自己的想法与…
主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
Wan: Open and Advanced Large-Scale Video Generative Models
[CVPR 2026]UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
[ECCV 2026 Oral] Implementation of "Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length"
Build Real-Time Knowledge Graphs for AI Agents
微舆:人人可用的多Agent舆情分析助手,打破信息茧房,还原舆情原貌,预测未来走向,辅助决策!从0实现,不依赖任何框架。