Skip to content
View nbasyl's full-sized avatar
🦇
I am Groot
🦇
I am Groot

Highlights

  • Pro

Block or report nbasyl

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

This code implements the algorithm of FIPO, a value-free RL recipe for eliciting deeper reasoning from a clean base model.

Python 130 6 Updated Apr 7, 2026
Python 18 3 Updated Jan 20, 2026

The official GitHub repo for the survey paper "A Survey on Diffusion Language Models".

1,174 59 Updated May 29, 2026

slime is an LLM post-training framework for RL Scaling.

Python 7,914 1,137 Updated Aug 14, 2026

Official implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Python 499 35 Updated May 20, 2026

Build compute kernels and load them from the Hub.

Python 723 119 Updated Aug 13, 2026

Scalable toolkit for efficient model reinforcement

Python 1,904 513 Updated Aug 14, 2026

Evaluate and improve models and agents using environments

Python 1,112 267 Updated Aug 14, 2026
Python 513 37 Updated Oct 16, 2025

✨✨Latest Papers and Benchmarks in Reasoning with Foundation Models

655 61 Updated Jun 16, 2025

[ICLR2026] Laser: Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Python 68 3 Updated May 22, 2025
Python 46 2 Updated Sep 27, 2025

[NeurIPS 2025] VeriThinker: Learning to Verify Makes Reasoning Model Efficient

Python 67 1 Updated Sep 27, 2025

[TMLR 2025] Efficient Reasoning Models: A Survey

Python 317 23 Updated Jun 26, 2026

Fully open data curation for reasoning models

Python 2,315 193 Updated Dec 2, 2025

Attribute statements generated by LLMs to preceding tokens using attention weights.

Jupyter Notebook 28 1 Updated Apr 22, 2025
Python 42 3 Updated Mar 26, 2025

s1: Simple test-time scaling

Python 6,663 756 Updated Jun 25, 2025

Work in progress.

Jupyter Notebook 81 10 Updated Nov 25, 2025

[ICLRW'26] EoRA: Fine-tuning-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation

Python 50 7 Updated Apr 21, 2026

Simple RL training for reasoning

Python 3,869 285 Updated Dec 23, 2025

Train transformer language models with reinforcement learning.

Python 19,070 2,906 Updated Aug 14, 2026

Fully open reproduction of DeepSeek-R1

Python 26,433 2,446 Updated Apr 2, 2026

YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open

Python 6,376 756 Updated Jun 4, 2025

A suite of image and video neural tokenizers

Jupyter Notebook 1,733 92 Updated Feb 11, 2025

LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

Python 1,226 200 Updated Aug 13, 2026

PyTorch native post-training library

Python 5,799 743 Updated Aug 13, 2026

Code for the ICLR 2023 paper "GPTQ: Accurate Post-training Quantization of Generative Pretrained Transformers".

Python 2,351 206 Updated Mar 27, 2024

Solutions to all questions of the book Introduction to the Theory of Computation, 3rd edition by Michael Sipser

1,846 290 Updated Dec 8, 2020

Solve puzzles. Learn CUDA.

Jupyter Notebook 12,401 944 Updated Sep 1, 2024
Next