Skip to content
View nbasyl's full-sized avatar
🦇
I am Groot
🦇
I am Groot

Highlights

  • Pro

Block or report nbasyl

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

This code implements the algorithm of FIPO, a value-free RL recipe for eliciting deeper reasoning from a clean base model.

Python 130 6 Updated Apr 7, 2026
Python 18 3 Updated Jan 20, 2026

The official GitHub repo for the survey paper "A Survey on Diffusion Language Models".

1,179 59 Updated Aug 17, 2026

slime is an LLM post-training framework for RL Scaling.

Python 8,145 1,163 Updated Aug 16, 2026

Official implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Python 500 36 Updated May 20, 2026

Build compute kernels and load them from the Hub.

Python 723 122 Updated Aug 19, 2026

Scalable toolkit for efficient model reinforcement

Python 1,921 521 Updated Aug 19, 2026

Evaluate and improve models and agents using environments

Python 1,127 278 Updated Aug 19, 2026
Python 514 37 Updated Oct 16, 2025

✨✨Latest Papers and Benchmarks in Reasoning with Foundation Models

655 61 Updated Jun 16, 2025

[ICLR2026] Laser: Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Python 68 3 Updated May 22, 2025
Python 46 2 Updated Sep 27, 2025

[NeurIPS 2025] VeriThinker: Learning to Verify Makes Reasoning Model Efficient

Python 67 1 Updated Sep 27, 2025

[TMLR 2025] Efficient Reasoning Models: A Survey

Python 318 23 Updated Jun 26, 2026

Fully open data curation for reasoning models

Python 2,318 194 Updated Dec 2, 2025

Attribute statements generated by LLMs to preceding tokens using attention weights.

Jupyter Notebook 28 1 Updated Apr 22, 2025
Python 42 3 Updated Mar 26, 2025

s1: Simple test-time scaling

Python 6,665 757 Updated Jun 25, 2025

Work in progress.

Jupyter Notebook 81 10 Updated Nov 25, 2025

[ICLRW'26] EoRA: Fine-tuning-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation

Python 50 8 Updated Apr 21, 2026

Simple RL training for reasoning

Python 3,873 284 Updated Dec 23, 2025

Train transformer language models with reinforcement learning.

Python 19,109 2,916 Updated Aug 19, 2026

Fully open reproduction of DeepSeek-R1

Python 26,437 2,446 Updated Apr 2, 2026

YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open

Python 6,395 758 Updated Jun 4, 2025

A suite of image and video neural tokenizers

Jupyter Notebook 1,733 93 Updated Feb 11, 2025

LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

Python 1,235 202 Updated Aug 19, 2026

PyTorch native post-training library

Python 5,799 745 Updated Aug 18, 2026

Code for the ICLR 2023 paper "GPTQ: Accurate Post-training Quantization of Generative Pretrained Transformers".

Python 2,355 207 Updated Mar 27, 2024

Solutions to all questions of the book Introduction to the Theory of Computation, 3rd edition by Michael Sipser

1,847 291 Updated Dec 8, 2020

Solve puzzles. Learn CUDA.

Jupyter Notebook 12,415 943 Updated Sep 1, 2024
Next