Lists (1)
Sort Name ascending (A-Z)
Starred repositories
🤖 The analysis of Claude Code
🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
An educational resource to help anyone learn deep reinforcement learning.
The corresponding codes and dataset for OneSearch series
OpenClaw-RL: Train any agent simply by talking
A Claude Code skill that acts as your daily 军师 (strategic research advisor).
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
Implement a reasoning LLM in PyTorch from scratch, step by step
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
AI agents running research on single-GPU nanochat training automatically
Open-source RL Framework with Online Teacher-Student Distillation
Pure Triton kernels for Qwen3.5-27B inference on NVIDIA B200
Sparse Transition Matrix-Accelerated Trie Index for Constrained Decoding (https://arxiv.org/abs/2602.22647)
Minimalistic 4D-parallelism distributed training framework for education purpose
Fast, small, and fully autonomous AI personal assistant infrastructure, any OS, any platform — deploy anywhere, swap anything 🦀
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
The official code for paper "Token-Level Collaborative Alignment for LLM-based Generative Recommendation"
[TMLR 26]: "UniRec: Unified Multimodal Encoding for LLM-Based Recommendations", Zijie Lei, Tao Feng, Zhigang Hua, Yan Xie, Guanyu Lin, Shuang Yang, Ge Liu, Jiaxuan You
"Unleashing the Potential of Sparse Attention on Long-term Behaviors for CTR Prediction." In Proceedings of WWW '26.
Our first fully AI generated deep learning system
Algorithm powering the For You feed on X
High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)
An Open Foundation Model and Benchmark to Accelerate Generative Recommendation
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
The official implementation of "ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning"