Skip to content
View ShuoZheLi's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report ShuoZheLi

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results
Python 6 Updated Oct 30, 2025
Python 2 Updated Apr 26, 2026

X-IL: Exploring the Design Space of Imitation Learning Policies

Python 63 5 Updated Mar 7, 2025

dLLM: Simple Diffusion Language Modeling

Python 2,692 281 Updated Jul 17, 2026

This project aims to collect the latest "call for reviewers" links from various top CS/ML/AI conferences/journals

1,178 50 Updated Feb 6, 2026

Minimal reproduction of DeepSeek R1-Zero

Python 13,242 1,575 Updated Feb 27, 2026

[NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards

Python 1,512 130 Updated Apr 17, 2026

A beautiful, simple, clean, and responsive Jekyll theme for academics

HTML 16,178 13,083 Updated Sep 21, 2026

Low ReSource Reinforcement Learning with CPU Offloading Training Support

Python 89 8 Updated Dec 10, 2025

One file implementation for LLM pretrain

Python 1 Updated Nov 10, 2025

Minecraft AI with LLMs+Mineflayer

JavaScript 5,779 929 Updated Jun 10, 2026

A benchmark environment for fully cooperative human-AI performance.

Jupyter Notebook 1,007 227 Updated Mar 22, 2025
JavaScript 39 8 Updated Jan 14, 2026
Python 1 Updated Jun 20, 2023

🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning

Python 27,716 5,720 Updated Sep 23, 2026
Python 1,312 134 Updated May 20, 2026

Build your own visual reasoning model

Jupyter Notebook 421 28 Updated Jan 13, 2026
Python 17 2 Updated Jun 25, 2025

[AAAI 2026] - Official repo for paper: "Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't"

Python 294 30 Updated Mar 11, 2026

Democratizing Reinforcement Learning for LLMs

Python 5,832 620 Updated Sep 12, 2026

RENT (Reinforcement Learning via Entropy Minimization) is an unsupervised method for training reasoning LLMs.

Python 42 7 Updated Oct 31, 2025

[ICLR2025 Spotlight] Advantage-Guided Distillation for Preference Alignment in Small Language Models

Python 27 1 Updated Feb 10, 2025

This repository collects papers for "A Survey on Knowledge Distillation of Large Language Models". We break down KD into Knowledge Elicitation and Distillation Algorithms, and explore the Skill & V…

1,314 75 Updated Mar 9, 2025
Jupyter Notebook 230 14 Updated Dec 23, 2025
Python 275 11 Updated May 14, 2025

[TMLR 2025] Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

791 42 Updated Feb 28, 2026
Next