Highlights
- Pro
Stars
Dataset of hackable TerminalBench-style tasks and exploit trajectories
[ICML 2026] Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
🔥 LLM-powered GPU kernel synthesis: Train models to convert PyTorch ops into optimized Triton kernels via SFT+RL. Multi-turn compilation feedback, cross-platform NVIDIA/AMD, Kernelbook + KernelBench
slime is an LLM post-training framework for RL Scaling.
A version of verl to support diverse tool use [TMLR 2026]
[COLM 2025] Official repository for R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agents
[EMNLP 2025] LightThinker: Thinking Step-by-Step Compression
[ICLR 2025] DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference
Monstertail / DeFT
Forked from LINs-lab/DeFT[ICLR 2025] DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning & ReCall: Learning to Reason with Tool Call for LLMs via Reinforcement Learning
[ACL2026] "MiniRAG: Making RAG Simpler with Small and Open-Sourced Language Models"
Based on the R1-Zero method, using rule-based rewards and GRPO on the Code Contests dataset.
Docker image NVIDIA GH200 machines - optimized for vllm serving and hf trainer finetuning
Moatless Testbeds allows you to create isolated testbed environments in a Kubernetes cluster where you can apply code changes through git patches and run tests or SWE-Bench evaluations.
Minimal reproduction of DeepSeek R1-Zero
A minimal language for Isabelle/HOL, designed for easing machine learning.
🤗 smolagents: a barebones library for agents that think in code.
Code to compute AnthroScore, a computational linguistic measure of anthropomorphism in text
Super-Efficient RLHF Training of LLMs with Parameter Reallocation
Formatron empowers everyone to control the format of language models' output with minimal overhead.