Skip to content
View gpt-dance's full-sized avatar
  • Peking University
  • BeiJing China

Block or report gpt-dance

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results
Python 170 6 Updated Jul 17, 2026

DCPO: Dynamic Adaptive Clipping for RL

Python 49 3 Updated Apr 1, 2026

Intentional Updates for Streaming Reinforcement Learning

Python 36 3 Updated Jun 6, 2026

🎓 系统性大语言模型构建课程|🛠️ 覆盖预训练数据工程、Tokenizer、Transformer、MoE、GPU 编程 (CUDA/Triton)、分布式训练、Scaling Laws、推理优化及对齐 (SFT/RLHF/GRPO)|🚀 6 个渐进式作业 + 代码驱动,建立 LLM 全栈认知体系

Jupyter Notebook 1,066 112 Updated Jul 23, 2026

ACL 2026 Main

Python 5 1 Updated Apr 14, 2026

A self-hosted ML coding practice platform. 68 problems from ReLU to flow matching — attention, training, RLHF, diffusion, and more. Instant feedback in the browser.

Python 1,199 111 Updated May 12, 2026

Vero: An Open RL Recipe for General Visual Reasoning

Python 134 11 Updated Jun 19, 2026
Python 735 44 Updated Mar 26, 2026

RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios

Python 591 59 Updated Jun 12, 2026

[ACL 2026] "OpenPhone: Mobile Agentic Foundation Models for AI Phone"

Python 927 182 Updated Jul 14, 2026

DART-GUI: Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation

Python 94 6 Updated Feb 26, 2026

[ICML 2026 Spotlight] On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models

Python 162 15 Updated Jun 8, 2026
Python 21 1 Updated Dec 3, 2025

[CVPR 2026] OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe

Python 164 5 Updated Mar 30, 2026
Jupyter Notebook 219 15 Updated Nov 25, 2025

DigiData: Training and evaluating general-purpose mobile control agents

Python 11 1 Updated Mar 4, 2026

Tiny-FSDP, a minimalistic re-implementation of the PyTorch FSDP

Python 112 9 Updated Aug 20, 2025

A curated list of RL resources

55 9 Updated Jun 11, 2026

从零构建大模型:从预训练到RLHF的完整实践

Python 2,678 209 Updated May 20, 2026

Train a 1B LLM with 1T tokens from scratch by personal

Jupyter Notebook 810 81 Updated Apr 27, 2025

[COLM 2025] Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources

Python 314 16 Updated Aug 25, 2025

Tiny-DeepSpeed, a minimalistic re-implementation of the DeepSpeed library

Python 53 13 Updated Aug 20, 2025

A lightweight reinforcement learning framework that integrates seamlessly into your codebase, empowering developers to focus on algorithms with minimal intrusion.

Python 106 2 Updated Aug 25, 2025

[AAAI 2026] GUI-G²: Gaussian Reward Modeling for GUI Grounding

Python 310 10 Updated Apr 15, 2026
Python 1,298 134 Updated May 20, 2026

Scaling Preference Data Curation via Human-AI Synergy

152 5 Updated Jul 3, 2025

slime is an LLM post-training framework for RL Scaling.

Python 7,616 1,093 Updated Jul 24, 2026

A simple framework to pre-train and fine-tune T5 model with pytorch-lightning and transformers

Python 8 Updated Apr 13, 2022

The AI developer platform. Use Weights & Biases to train and fine-tune models, and manage models from experimentation to production.

Python 11,202 883 Updated Jul 24, 2026
Next