Skip to content
View BetuBin18070's full-sized avatar
  • University of Chinese Academy of Sciences
  • Beijing, China
  • 00:24 (UTC -12:00)

Block or report BetuBin18070

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

The Entropy Mechanism of Reinforcement Learning for Large Language Model Reasoning.

Python 449 15 Updated Jul 11, 2025

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.

Jupyter Notebook 19,787 1,829 Updated Jan 30, 2026

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

Python 74,093 9,067 Updated Aug 13, 2026

A collection of LLM papers, blogs, and projects, with a focus on OpenAI o1 🍓 and reasoning techniques.

6,899 369 Updated Dec 17, 2025

EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL

Python 5,115 383 Updated Jul 30, 2026
Python 80 4 Updated Nov 19, 2024

Minimal reproduction of DeepSeek R1-Zero

Python 13,227 1,578 Updated Feb 27, 2026

An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

Python 9,911 999 Updated Aug 13, 2026

R1-onevision, a visual language model capable of deep CoT reasoning.

Python 582 16 Updated Apr 13, 2025

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,958 4,403 Updated Aug 14, 2026

Solve Visual Understanding with Reinforced VLMs

Python 6,016 385 Updated Jul 7, 2026

Train transformer language models with reinforcement learning.

Python 19,069 2,906 Updated Aug 14, 2026

Official Repo for Open-Reasoner-Zero

Python 2,098 120 Updated Jun 2, 2025

This repository collects various works that reproduce DeepSeek R1, as well as works related to DeepSeek R1 and the DeepSeek series.

19 Updated Apr 27, 2025

Implementations and examples of common offline policy evaluation methods in Python.

Python 220 25 Updated Feb 11, 2023

Code for the paper "Training Diffusion Models with Reinforcement Learning"

Python 578 34 Updated Jul 5, 2023

Simple RL training for reasoning

Python 3,869 285 Updated Dec 23, 2025

Fully open reproduction of DeepSeek-R1

Python 26,433 2,446 Updated Apr 2, 2026

Solutions of Reinforcement Learning, An Introduction

Jupyter Notebook 2,428 510 Updated Jul 10, 2025

Python Implementation of Reinforcement Learning: An Introduction

Python 14,747 4,961 Updated Aug 9, 2024

Solutions to exercises in Reinforcement Learning: An Introduction (2nd Edition).

Jupyter Notebook 412 82 Updated Jul 24, 2023

The official repository for NeurIPS 2024 Oral <Maximum Entropy Inverse Reinforcement Learning of Diffusion Models with Energy-Based Models>

Python 33 5 Updated Mar 20, 2025

Simulation platform for general-purpose robotics & embodied AI learning.

Python 29,741 2,835 Updated Aug 13, 2026

Repository of notes, code and notebooks in Python for the book Pattern Recognition and Machine Learning by Christopher Bishop

Jupyter Notebook 2,626 544 Updated Jul 25, 2022

Trustworthy AI related projects

Python 1,133 249 Updated Jun 1, 2026

The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery 🧑‍🔬

Jupyter Notebook 14,392 2,038 Updated Dec 19, 2025

🌴爱情树,将相爱的时刻永远珍藏 (微信,QQ可完美查看)https://ajlovechina.github.io/LoveTree/

JavaScript 422 907 Updated Aug 7, 2024
Next