Stars
🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!
Industrial audio online policy distillation (OPD) training stack for ASR and TTS, distilling compact audio models from stronger teacher models.
Research of DeepSeek Engram Architecture based on Qwen-3 and Stable Diffusion series.
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
[ICCV 2025] Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
Understand Human Behavior to Align True Needs
StreamDiffusion: A Pipeline-Level Solution for Real-Time Interactive Generation
Official implementation of "MoMask: Generative Masked Modeling of 3D Human Motions (CVPR2024)"
[ICCV2023] Delicate Textured Mesh Recovery from NeRF via Adaptive Surface Refinement
A Unified Framework for Surface Reconstruction
AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
Erasing Appearance Preservation in Optimization-based Smoothing (ECCV 2020)
A PyTorch3D walkthrough and a Medium article 👋 on how to render 3D .obj meshes from various viewpoints to create 2D images.