Skip to content
View Ki6an's full-sized avatar
👾
👾

Block or report Ki6an

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Stars

training

21 repositories

USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference

Python 685 82 Updated May 21, 2026

OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Python 1,852 132 Updated Jan 17, 2025

DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.

Python 42,930 4,929 Updated Aug 14, 2026
Jupyter Notebook 4 1 Updated Feb 16, 2024

Source code of our EMNLP 2024 paper "FactAlign: Long-form Factuality Alignment of Large Language Models"

Jupyter Notebook 19 1 Updated Oct 3, 2024

Entropy Based Sampling and Parallel CoT Decoding

Python 3,431 317 Updated Nov 13, 2024

[NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward

Python 958 78 Updated Feb 16, 2025

Code and Configs for Asynchronous RLHF: Faster and More Efficient RL for Language Models

Python 68 11 Updated Mar 5, 2026

Democratizing Reinforcement Learning for LLMs

Python 5,785 605 Updated Aug 14, 2026

High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)

Python 10,265 1,153 Updated Apr 20, 2026

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …

Python 15,156 1,591 Updated Aug 13, 2026

AllenAI's post-training codebase

Python 3,831 573 Updated Aug 14, 2026

SkyRL: A Modular Full-stack RL Library for LLMs

Python 2,151 403 Updated Aug 14, 2026

A scalable automated alignment method for large language models. Resources for "Aligning Large Language Models via Self-Steering Optimization".

Python 20 3 Updated Nov 21, 2024

Agentic RL Training at Scale

Python 1,912 397 Updated Aug 14, 2026

Less is More: Task-aware Layer-wise Distillation for Language Model Compression (ICML2023)

Python 40 5 Updated Aug 28, 2023

The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.

Python 5,665 574 Updated Aug 14, 2026

[ICLR 2025] COAT: Compressing Optimizer States and Activation for Memory-Efficient FP8 Training

Python 263 27 Updated Aug 9, 2025

PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署)

C++ 24,051 6,014 Updated Aug 14, 2026

RL environments + evals for AI agents. Define once, train anything.

Python 291 68 Updated Aug 14, 2026

The Amazon S3 Connector for PyTorch delivers high throughput for PyTorch training jobs that access and store data in Amazon S3.

Python 213 35 Updated Jul 27, 2026