Skip to content
View Shenzhi-Wang's full-sized avatar

Block or report Shenzhi-Wang

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. hiyouga/EasyR1 hiyouga/EasyR1 Public

    EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL

    Python 5.2k 397

  2. LeapLabTHU/cooragent LeapLabTHU/cooragent Public

    Official Repository of Cooragent.

    Python 1.7k 143

  3. Llama3-Chinese-Chat Llama3-Chinese-Chat Public

    This is the first Chinese chat model specifically fine-tuned for Chinese through ORPO based on the Meta-Llama-3-8B-Instruct model.

    316 21

  4. Beyond-the-80-20-Rule-RLVR Beyond-the-80-20-Rule-RLVR Public

    The open-source code for the NeurIPS 2025 paper, "Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning."

    Python 68 4

  5. LeapLabTHU/FamO2O LeapLabTHU/FamO2O Public

    Repository of "Train Once, Get a Family: State-Adaptive Balances for Offline-to-Online Reinforcement Learning" (NeurIPS 2023 Spotlight)

    Python 41 2

  6. recon recon Public

    The official source code for "Boosting LLM Agents with Recursive Contemplation for Effective Deception Handling" (ACL 2024, Findings)

    Python 15