-
@NVlabs, HKUST
- ?
- https://nbasyl.github.io
Highlights
- Pro
Stars
This code implements the algorithm of FIPO, a value-free RL recipe for eliciting deeper reasoning from a clean base model.
The official GitHub repo for the survey paper "A Survey on Diffusion Language Models".
slime is an LLM post-training framework for RL Scaling.
Official implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Build compute kernels and load them from the Hub.
Scalable toolkit for efficient model reinforcement
Evaluate and improve models and agents using environments
✨✨Latest Papers and Benchmarks in Reasoning with Foundation Models
[ICLR2026] Laser: Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
[NeurIPS 2025] VeriThinker: Learning to Verify Makes Reasoning Model Efficient
[TMLR 2025] Efficient Reasoning Models: A Survey
Fully open data curation for reasoning models
Attribute statements generated by LLMs to preceding tokens using attention weights.
[ICLRW'26] EoRA: Fine-tuning-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation
Train transformer language models with reinforcement learning.
Fully open reproduction of DeepSeek-R1
YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open
A suite of image and video neural tokenizers
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Code for the ICLR 2023 paper "GPTQ: Accurate Post-training Quantization of Generative Pretrained Transformers".
Solutions to all questions of the book Introduction to the Theory of Computation, 3rd edition by Michael Sipser