training
USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference
OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Source code of our EMNLP 2024 paper "FactAlign: Long-form Factuality Alignment of Large Language Models"
Entropy Based Sampling and Parallel CoT Decoding
[NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward
Code and Configs for Asynchronous RLHF: Faster and More Efficient RL for Language Models
High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
SkyRL: A Modular Full-stack RL Library for LLMs
A scalable automated alignment method for large language models. Resources for "Aligning Large Language Models via Self-Steering Optimization".
Less is More: Task-aware Layer-wise Distillation for Language Model Compression (ICML2023)
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
[ICLR 2025] COAT: Compressing Optimizer States and Activation for Memory-Efficient FP8 Training
PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署)
RL environments + evals for AI agents. Define once, train anything.
The Amazon S3 Connector for PyTorch delivers high throughput for PyTorch training jobs that access and store data in Amazon S3.