Skip to content
View mukhal's full-sized avatar

Block or report mukhal

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

Python 78 5 Updated Jul 25, 2026

Lean math proofs generated by AlphaProof Nexus and accompanying natural language prose proofs.

Lean 280 20 Updated Jul 21, 2026

Convert any Repo into an RL Environment

Python 469 72 Updated Jul 23, 2026

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

Python 1,789 325 Updated Jul 26, 2026

A curated list of papers and resources on Reward Hacking, Emergent Misalignment, and Proxy Exploitation in Large Models

42 4 Updated Apr 17, 2026

🔥 LeetCode for PyTorch — practice implementing softmax, attention, GPT-2 and more from scratch with instant auto-grading. Jupyter-based, self-hosted or try online.

Jupyter Notebook 4,400 384 Updated May 25, 2026

A testbed for studying the emergence and generalization of reward hacking

Python 9 4 Updated Mar 10, 2026

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

Python 4,630 755 Updated May 17, 2026

generate coding exercises from any github repo

Python 2 Updated Oct 28, 2025

Provide with pre-build flash-attention 2 and 3 package wheels on Linux and Windows using GitHub Actions

Python 1,638 74 Updated Jul 26, 2026

Processed / Cleaned Data for Paper Copilot

Python 952 47 Updated Jul 1, 2026
Python 1,298 135 Updated May 20, 2026

Eliciting Long CoT from a Short CoT Model

Python 8 Updated May 16, 2025

Single File, Single GPU, From Scratch, Efficient, Full Parameter Tuning library for "RL for LLMs"

Jupyter Notebook 626 57 Updated Oct 7, 2025

A comprehensive collection of process reward models.

176 4 Updated Jun 6, 2026

[TMLR] Process Reward Models That Think

Python 89 8 Updated Nov 29, 2025
Python 10 11 Updated Nov 14, 2025

[NeurIPS'21 Outstanding Paper] Library for reliable evaluation on RL and ML benchmarks, even with only a handful of seeds.

Jupyter Notebook 880 49 Updated Aug 12, 2024

s1: Simple test-time scaling

Python 6,660 757 Updated Jun 25, 2025

Our library for RL environments + evals

Python 4,400 612 Updated Jul 25, 2026

LLM-Merging: Building LLMs Efficiently through Merging

Jupyter Notebook 208 44 Updated Sep 24, 2024

pipreqs - Generate pip requirements.txt file based on imports of any project. Looking for maintainers to move this project forward.

Python 7,462 420 Updated Mar 30, 2026
JavaScript 4,293 1,942 Updated Jun 21, 2024

Official repository for ACL 2025 paper "ProcessBench: Identifying Process Errors in Mathematical Reasoning"

Python 190 18 Updated May 20, 2025

Recipes to scale inference-time compute of open models

Python 1,130 131 Updated May 26, 2026

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities. ACM Computing Surveys, 2026.

769 47 Updated Jul 17, 2026

Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline

Python 878 83 Updated Jul 20, 2026

A framework for the evaluation of autoregressive code generation language models.

Python 1,055 263 Updated Jul 22, 2025
Next