Stars
A Framework for LLM-based Multi-Agent Reinforced Training and Inference
✨✨Latest Advances on Multimodal Large Language Models
RAGEN leverages reinforcement learning to train LLM reasoning agents in interactive, stochastic environments.
Samples for CUDA Developers which demonstrates features in CUDA Toolkit
📚LeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginners🐑, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.🎉
基于通义千问 Qwen2.5-Omni 的实时语音对话系统,使用在线API服务,支持实时语音交互、动态语音活动检测和流式音频处理。A real-time voice conversation system based on Qwen2.5-Omni Online-API, supporting real-time voice interaction, dynamic voice activi…
✅(已完结)超级全面的 深度学习 笔记【土堆 Pytorch】【李沐 动手学深度学习】【吴恩达 深度学习】【大飞 大模型Agent】
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
✨终生持续更新✨ 计算机基础自学笔记/心得/实验/资源汇总;课程:数据结构、操作系统(MIT6.S081)、分布式系统(MIT6.824)等
Using ChatGPT to connect with FreeSWITCH, creating an intelligent phone robot.
official implementation of ICLR'2025 paper: Rethinking Bradley-Terry Models in Preference-based Reward Modeling: Foundations, Theory, and Alternatives
简单实现VAD+声纹锁+SenseVoice完成类语音实时转录的小项目
PyTorch version of Stable Baselines, reliable implementations of reinforcement learning algorithms.
A library with extensible implementations of DPO, KTO, PPO, ORPO, and other human-aware loss functions (HALOs).
Implementations of selected inverse reinforcement learning algorithms.
This repository collects papers for "A Survey on Knowledge Distillation of Large Language Models". We break down KD into Knowledge Elicitation and Distillation Algorithms, and explore the Skill & V…
每个人都能看懂的大模型知识分享,LLMs春/秋招大模型面试前必看,让你和面试官侃侃而谈
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
My course work solutions and quiz answers
Train transformer language models with reinforcement learning.
ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search (NeurIPS 2024)