-
NVIDIA
- New York
- http://mukhal.github.io
- @mkhalifaaaa
Stars
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
Lean math proofs generated by AlphaProof Nexus and accompanying natural language prose proofs.
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
A curated list of papers and resources on Reward Hacking, Emergent Misalignment, and Proxy Exploitation in Large Models
🔥 LeetCode for PyTorch — practice implementing softmax, attention, GPT-2 and more from scratch with instant auto-grading. Jupyter-based, self-hosted or try online.
A testbed for studying the emergence and generalization of reward hacking
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
generate coding exercises from any github repo
Provide with pre-build flash-attention 2 and 3 package wheels on Linux and Windows using GitHub Actions
Processed / Cleaned Data for Paper Copilot
yunx-z / ThinkLogit
Forked from alisawuffles/proxy-tuningEliciting Long CoT from a Short CoT Model
Single File, Single GPU, From Scratch, Efficient, Full Parameter Tuning library for "RL for LLMs"
A comprehensive collection of process reward models.
[NeurIPS'21 Outstanding Paper] Library for reliable evaluation on RL and ML benchmarks, even with only a handful of seeds.
Our library for RL environments + evals
LLM-Merging: Building LLMs Efficiently through Merging
pipreqs - Generate pip requirements.txt file based on imports of any project. Looking for maintainers to move this project forward.
Official repository for ACL 2025 paper "ProcessBench: Identifying Process Errors in Mathematical Reasoning"
Recipes to scale inference-time compute of open models
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities. ACM Computing Surveys, 2026.
Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline
A framework for the evaluation of autoregressive code generation language models.