Skip to content
View Ashitaka2's full-sized avatar

Block or report Ashitaka2

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Design of noise, interpolation schedules, and sources in generative dynamics, with enhanced numerical performance

Jupyter Notebook 8 3 Updated May 16, 2026

The codebase of our paper "Improving the Training of Rectified Flows", NeurIPS 2024

Python 133 8 Updated Oct 18, 2024

State-of-the-Art Embeddings, Retrieval, and Reranking

Python 19,011 2,855 Updated Aug 14, 2026

Emergent Hierarchical Reasoning in LLMs/VLMs through Reinforcement Learning [ICLR26]

Python 64 2 Updated Apr 11, 2026

Code for Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities (NeurIPS'24)

Python 36 4 Updated Dec 17, 2024
Python 8 1 Updated Oct 10, 2024

a neuroscience based model simulating basal ganglia for mouse maze pathfinding with q-learning

Jupyter Notebook 3 2 Updated Nov 20, 2023

Minimal reproduction of DeepSeek R1-Zero

Python 13,226 1,578 Updated Feb 27, 2026

An Open-source RL System from ByteDance Seed and Tsinghua AIR

Python 1,855 85 Updated May 11, 2025

The Entropy Mechanism of Reinforcement Learning for Large Language Model Reasoning.

Python 449 15 Updated Jul 11, 2025

Train transformer language models with reinforcement learning.

Python 19,080 2,910 Updated Aug 15, 2026

[NeurIPS 2025] TTRL: Test-Time Reinforcement Learning

Python 1,110 82 Updated Apr 15, 2026
Python 361 20 Updated Jul 29, 2025

Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

Python 2,519 531 Updated Aug 11, 2026
Python 1,179 58 Updated Jan 10, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,974 4,410 Updated Aug 15, 2026
Python 554 65 Updated Jan 2, 2025

PyTorch implementation of Advantage Actor Critic (A2C), Proximal Policy Optimization (PPO), Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation (ACKT…

Python 3,902 840 Updated May 29, 2022

Code for NeurIPS'24 paper 'Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization'

Python 241 26 Updated Jul 19, 2025

Simple RL training for reasoning

Python 3,869 285 Updated Dec 23, 2025

Code for the paper "The Impact of Positional Encoding on Length Generalization in Transformers", NeurIPS 2023

Python 139 7 Updated Apr 30, 2024
Jupyter Notebook 246 75 Updated May 10, 2024

An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

Python 9,917 1,002 Updated Aug 13, 2026
Python 221 10 Updated Feb 20, 2025

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Jupyter Notebook 264 29 Updated May 14, 2025

s1: Simple test-time scaling

Python 6,663 757 Updated Jun 25, 2025

Cliff walking reinforcement learning example, with a variety of RL algorithms

Python 15 3 Updated Dec 5, 2023

LLM training in simple, raw C/CUDA

Cuda 30,813 3,734 Updated Jun 26, 2025

The simplest, fastest repository for training/finetuning medium-sized GPTs.

Python 62,136 10,709 Updated Nov 12, 2025
Next