-
University of Chinese Academy of Sciences
- Beijing, China
-
00:24
(UTC -12:00)
Stars
The Entropy Mechanism of Reinforcement Learning for Large Language Model Reasoning.
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
A collection of LLM papers, blogs, and projects, with a focus on OpenAI o1 🍓 and reasoning techniques.
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
Minimal reproduction of DeepSeek R1-Zero
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
R1-onevision, a visual language model capable of deep CoT reasoning.
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
Solve Visual Understanding with Reinforced VLMs
Train transformer language models with reinforcement learning.
Official Repo for Open-Reasoner-Zero
This repository collects various works that reproduce DeepSeek R1, as well as works related to DeepSeek R1 and the DeepSeek series.
Implementations and examples of common offline policy evaluation methods in Python.
Code for the paper "Training Diffusion Models with Reinforcement Learning"
Fully open reproduction of DeepSeek-R1
Solutions of Reinforcement Learning, An Introduction
Python Implementation of Reinforcement Learning: An Introduction
Solutions to exercises in Reinforcement Learning: An Introduction (2nd Edition).
The official repository for NeurIPS 2024 Oral <Maximum Entropy Inverse Reinforcement Learning of Diffusion Models with Energy-Based Models>
Simulation platform for general-purpose robotics & embodied AI learning.
Repository of notes, code and notebooks in Python for the book Pattern Recognition and Machine Learning by Christopher Bishop
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery 🧑🔬
AJLoveChina / LoveTree
Forked from lvshaoge/ebao.space🌴爱情树,将相爱的时刻永远珍藏 (微信,QQ可完美查看)https://ajlovechina.github.io/LoveTree/