Skip to content
View XL2248's full-sized avatar

Block or report XL2248

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Official repository for the paper "Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation"

Python 283 14 Updated May 28, 2026

Code for "Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement Learning".

Python 28 Updated Nov 11, 2025

UR2: Unify RAG and Reasoning through Reinforcement Learning

Python 131 11 Updated May 26, 2026

Learning to Generate STRUCTURED Output with Schema Reinforcement Learning

Python 26 6 Updated Mar 2, 2025
Shell 6 3 Updated Sep 10, 2025

Multilingual and Multiculture Benchmark and LLM

Python 43 1 Updated Jul 30, 2026

Official repository for DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression

Python 6 Updated Aug 23, 2025

Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning

Python 5 1 Updated Jan 12, 2026

Accepted by ACL 2025

Python 30 Updated Aug 13, 2025

The official repo of One RL to See Them All: Visual Triple Unified Reinforcement Learning

Python 332 17 Updated May 31, 2025

RAGEN leverages reinforcement learning to train LLM reasoning agents in interactive, stochastic environments.

Python 2,770 229 Updated Jul 24, 2026

Scalable RL solution for advanced reasoning of language models

Python 1,867 116 Updated Mar 18, 2025

Reproduce R1 Zero on Logic Puzzle

Python 2,449 163 Updated Mar 20, 2025

a benckmark for evaluating logical reasoning of LLMs

Python 23 4 Updated Jan 25, 2024

GAOGAO-Bench-Updates is a supplement to the GAOKAO-Bench, a dataset to evaluate large language models.

Python 48 5 Updated Jan 7, 2025

Democratizing Reinforcement Learning for LLMs

Python 5,784 606 Updated Aug 16, 2026

Deep Reasoning Translation (DRT) Project

242 9 Updated Sep 1, 2025

华中科学大学数学分析与高代代数考研真题

TeX 10 1 Updated Jun 10, 2023

中国科学院大学,601高等数学甲,历年考研真题收集整理

TeX 13 Updated Aug 4, 2025

LLM evaluation on 2024 Chinese Gaokao Mathematics — zero-contamination benchmark with dual prompt formats

21 2 Updated Apr 15, 2026

The most comprehensive database of Chinese poetry 🧶最全中华古诗词数据库, 唐宋两朝近一万四千古诗人, 接近5.5万首唐诗加26万宋诗. 两宋时期1564位词人,21050首词。

JavaScript 53,133 10,695 Updated Jun 17, 2026

[ICML 2024] Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning

Python 49 5 Updated May 12, 2024

A family of open-sourced Mixture-of-Experts (MoE) Large Language Models

Python 1,693 86 Updated Mar 8, 2024

Fast inference engine for Transformer models

C++ 4,620 513 Updated Aug 15, 2026

MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集。对标chatGPT训练的40T数据。MNBVC数据集不但包括主流文化,也包括各个小众文化甚至火星文的数据。MNBVC数据集包括新闻、作文、小说、书籍、杂志、论文、台词、帖子、wiki、古诗、歌词、商品介绍、笑话、糗事、聊天记录等一切形式的纯文本中文数据。

4,262 296 Updated Aug 15, 2026
Jsonnet 407 58 Updated May 2, 2024

This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.

Python 12,152 1,061 Updated Mar 8, 2026

✨✨Latest Advances on Multimodal Large Language Models

17,977 1,133 Updated Aug 14, 2026

Implementation of AudioLM, a SOTA Language Modeling Approach to Audio Generation out of Google Research, in Pytorch

Python 2,627 278 Updated Jan 12, 2025
Next