Skip to content
View okoge-kaz's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Organizations

@rioyokotalab

Block or report okoge-kaz

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models

Python 3,358 304 Updated Aug 11, 2026

MrlX: A Multi-Agent Reinforcement Learning Framework

Python 220 12 Updated Jan 19, 2026

Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework

Python 337 27 Updated Jan 17, 2026

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective

Python 2 Updated Jun 24, 2026

[ICML 2026] Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs

Python 33 3 Updated Apr 27, 2026

Framework for evaluating and improving agents

Python 4,098 1,516 Updated Aug 11, 2026

Measuring frontier coding agents on original, long-horizon engineering tasks

Python 1,345 87 Updated Aug 6, 2026

Our library for RL environments + evals

Python 4,490 638 Updated Aug 11, 2026

ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale

C++ 659 225 Updated Apr 25, 2026

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

Python 1,951 352 Updated Aug 11, 2026

Environments by the Prime Intellect Research Team

Python 115 36 Updated Aug 11, 2026

A list of cloud sandbox providers for AI agents. Information sourced exclusively from official docs and landing pages.

84 24 Updated Jul 6, 2026
Python 898 85 Updated Aug 6, 2026

Allow torch tensor memory to be released and resumed later

Python 267 69 Updated Aug 9, 2026

A fork of sgl-model-gateway for slime.

Rust 8 Updated Jul 3, 2026

Open Machine Learning Compiler Framework

Python 13,665 3,948 Updated Aug 11, 2026

Efficient Long-context Language Model Training by Core Attention Disaggregation

Python 106 7 Updated Apr 7, 2026
C++ 382 43 Updated Jan 28, 2026

An LLM post-training framework with vLLM for RL Scaling

Python 410 74 Updated Aug 3, 2026

Scalable RL for Any Agent and Sandbox.

Python 575 66 Updated Aug 7, 2026

[ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Python 404 26 Updated Mar 30, 2026

A lightweight inference engine supporting speculative speculative decoding (SSD).

Python 983 78 Updated May 10, 2026

[ICML 2026] d3LLM: Ultra-Fast Diffusion LLM 🚀

Python 149 10 Updated May 1, 2026

An interface library for RL post training with environments.

Python 2,492 423 Updated Aug 11, 2026

mKernel: fast multi-node, multi-GPU fused kernels

Cuda 264 24 Updated Aug 10, 2026

Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation

8,039 293 Updated May 15, 2025

Implementation of the sparse attention pattern proposed by the Deepseek team in their "Native Sparse Attention" paper

Python 811 53 Updated Aug 15, 2025

🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"

Python 1,017 53 Updated Feb 5, 2026

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++ 6,230 1,064 Updated Aug 11, 2026

A scalable asynchronous reinforcement learning implementation with in-flight weight updates.

Python 433 50 Updated Aug 5, 2026
Next