-
Institute of Science Tokyo
- Tokyo Japan
-
13:59
(UTC +09:00) - https://okoge-kaz.github.io/
- @kazukifujii
- in/kazuki-fujii
- https://scholar.google.co.jp/citations?user=jHXLs2wAAAAJ&hl=en
Highlights
- Pro
Stars
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
MrlX: A Multi-Agent Reinforcement Learning Framework
Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective
[ICML 2026] Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
Framework for evaluating and improving agents
Measuring frontier coding agents on original, long-horizon engineering tasks
Our library for RL environments + evals
ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
Environments by the Prime Intellect Research Team
A list of cloud sandbox providers for AI agents. Information sourced exclusively from official docs and landing pages.
Allow torch tensor memory to be released and resumed later
Efficient Long-context Language Model Training by Core Attention Disaggregation
An LLM post-training framework with vLLM for RL Scaling
[ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
A lightweight inference engine supporting speculative speculative decoding (SSD).
An interface library for RL post training with environments.
mKernel: fast multi-node, multi-GPU fused kernels
Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation
Implementation of the sparse attention pattern proposed by the Deepseek team in their "Native Sparse Attention" paper
🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
A scalable asynchronous reinforcement learning implementation with in-flight weight updates.