Starred repositories
Source code of paper: A Stronger Mixture of Low-Rank Experts for Fine-Tuning Foundation Models. (ICML 2025)
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
Digital Avatar Conversational System - Linly-Talker. 😄✨ Linly-Talker is an intelligent AI system that combines large language models (LLMs) with visual models to create a novel human-AI interaction…
Open-source framework for conversational voice AI agents
Real time web based Speech-to-Text app with Streamlit
open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.
Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
zero-shot voice conversion & singing voice conversion, with real-time support
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
Build local voice agents with open-source models
Train transformer language models with reinforcement learning.
Reference implementation for DPO (Direct Preference Optimization)
Noise reduction in python using spectral gating (speech, bioacoustics, audio, time-domain signals)
Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
Open-Sora: Democratizing Efficient Video Production for All
VideoSys: An easy and efficient system for video generation
This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
Official PyTorch Implementation of "Scalable Diffusion Models with Transformers"
Implementation of MagViT2 Tokenizer in Pytorch
OpenMMLab Pose Estimation Toolbox and Benchmark.
《Hello 算法》:动画图解、一键运行的数据结构与算法教程。支持简中、繁中、English、日本語,提供 Python, Java, C++, C, C#, JS, Go, Swift, Rust, Ruby, Kotlin, TS, Dart 等代码实现
Official Repo for the Paper: CHATANYTHING: FACETIME CHAT WITH LLM-ENHANCED PERSONAS
Pytorch official implementation for our paper "HyperLips: Hyper Control Lips with High Resolution Decoder for Talking Face Generation".
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
PyTorch Re-Implementation of "The Sparsely-Gated Mixture-of-Experts Layer" by Noam Shazeer et al. https://arxiv.org/abs/1701.06538
[CVPR'24 Highlight] Official PyTorch implementation of CoDeF: Content Deformation Fields for Temporally Consistent Video Processing