Skip to content
View zkx06111's full-sized avatar
🤔
Thinking about what to do.
🤔
Thinking about what to do.

Highlights

  • Pro

Block or report zkx06111

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Dataset of hackable TerminalBench-style tasks and exploit trajectories

Python 35 2 Updated Apr 18, 2026

[ICML 2026] Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks

Python 48 7 Updated Aug 16, 2026

🔥 LLM-powered GPU kernel synthesis: Train models to convert PyTorch ops into optimized Triton kernels via SFT+RL. Multi-turn compilation feedback, cross-platform NVIDIA/AMD, Kernelbook + KernelBench

Python 149 6 Updated Nov 10, 2025

slime is an LLM post-training framework for RL Scaling.

Python 8,093 1,152 Updated Aug 16, 2026
Python 17 Updated Jan 27, 2026

A version of verl to support diverse tool use [TMLR 2026]

Python 1,031 89 Updated Jul 15, 2026

[COLM 2025] Official repository for R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agents

Python 320 63 Updated Jul 13, 2025

[EMNLP 2025] LightThinker: Thinking Step-by-Step Compression

Python 165 6 Updated Jun 22, 2026
Python 37 2 Updated Mar 19, 2025
Jupyter Notebook 25 3 Updated May 27, 2026

[ICLR 2025] DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference

Jupyter Notebook 54 2 Updated Jun 17, 2025

[ICLR 2025] DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference

Jupyter Notebook 2 Updated Mar 5, 2025

ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning & ReCall: Learning to Reason with Tool Call for LLMs via Reinforcement Learning

Python 1,428 89 Updated May 16, 2025

[ACL2026] "MiniRAG: Making RAG Simpler with Small and Open-Sourced Language Models"

Python 2,007 257 Updated Oct 16, 2025

Reproducing R1 for Code with Reliable Rewards

Python 317 20 Updated May 5, 2025
Python 139 12 Updated Mar 20, 2025

Based on the R1-Zero method, using rule-based rewards and GRPO on the Code Contests dataset.

Python 18 3 Updated Apr 22, 2025

Docker image NVIDIA GH200 machines - optimized for vllm serving and hf trainer finetuning

Python 56 6 Updated Jun 22, 2026
Python 336 18 Updated May 31, 2025

RL Scaling and Test-Time Scaling (ICML'25)

116 1 Updated Jan 23, 2025

Moatless Testbeds allows you to create isolated testbed environments in a Kubernetes cluster where you can apply code changes through git patches and run tests or SWE-Bench evaluations.

Python 14 3 Updated Apr 9, 2025

Minimal reproduction of DeepSeek R1-Zero

Python 13,224 1,578 Updated Feb 27, 2026

A minimal language for Isabelle/HOL, designed for easing machine learning.

Python 31 Updated Aug 14, 2026

🤗 smolagents: a barebones library for agents that think in code.

Python 28,841 2,865 Updated Jul 21, 2026

珠算代码大模型(Abacus Code LLM)

57 2 Updated Sep 26, 2024

Code to compute AnthroScore, a computational linguistic measure of anthropomorphism in text

Python 19 2 Updated Mar 31, 2025

AllenAI's post-training codebase

Python 3,832 577 Updated Aug 16, 2026

Super-Efficient RLHF Training of LLMs with Parameter Reallocation

Python 336 22 Updated Apr 24, 2025

Formatron empowers everyone to control the format of language models' output with minimal overhead.

Python 237 9 Updated Jun 7, 2025
Next