-
Peking University
- Beijing
Stars
AxisRL is an agentic RL post-training framework built on SGLang rollout, Megatron training, and real-world agent workflows.
AxisAgentic: An Extensible Runtime and Trajectory-Collection Framework for Long-Horizon Agents.
A Claude Code plugin that shows what's happening - context usage, active tools, running agents, and todo progress
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
🏆 Top-1 on 5+ benchmarks | Web UI | Supports MiroThinker, Claude, Kimi, OpenAI
OpenSeeker: A search agent with open-source data and models
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74.0 and 75.3 on the BrowseComp and BrowseComp Zh, respectively.
[TMLR 2024] Official implementation of "NuTime: Numerically Multi-Scaled Embedding for Large-Scale Time-Series Pretraining".
An elegant \LaTeX\ résumé template. 大陆镜像 https://gods.coding.net/p/resume/git
A Survey of Reinforcement Learning for Large Reasoning Models
The official Python library for the OpenAI API
gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
Official repository for the paper "LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code"
Research and development (R&D) is crucial for the enhancement of industrial productivity, especially in the AI era, where the core aspects of R&D are mainly focused on data and models. We are commi…
An Open-source RL System from ByteDance Seed and Tsinghua AIR
The rule-based evaluation subset and code implementation of Omni-MATH
repo for paper https://arxiv.org/abs/2504.13837
[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Fast and memory-efficient exact attention
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Fully open reproduction of DeepSeek-R1
SGLang is a high-performance serving framework for large language models and multimodal models.
Unleashing the Power of Reinforcement Learning for Math and Code Reasoners