Stars
Welcome to KernelBench-Verified. This repository provides a robust, realistic evaluation framework for assessing custom CUDA kernels generated by Large Language Models (LLMs).
[ICML 2026] AdaMEM: Test-Time Adaptive Memory for Language Agents
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
A single-file implementation of KV cache paged attention
mini-swe-agent-plus: a tiny (~100 LOC) GitHub issue fixer—now with a robust multi-line text edit tool.
A Practitioner's Guide to M(eow)ti Turn Agentic ReinfOrcement learning
[NeurIPS 2025 D&B Track] MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
The official implementation of "ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering"
About The official GitHub page for ''Unlocking General Long Chain-of-Thought Reasoning Capabilities of Large Language Models via Representation Engineering'' Resources
MLGym A New Framework and Benchmark for Advancing AI Research Agents
SGLang is a high-performance serving framework for large language models and multimodal models.
This repository contains the 1st place source code of Track II: Backdoor Trigger Recovery for Models in The Competition for LLM and Agent Safety 2024 at NeurIPS 2024.
A curated list of papers on LLMs and agents for scientific research and development
Official code for paper: Chain of Ideas: Revolutionizing Research via Novel Idea Development with LLM Agents
DSIR large-scale data selection framework for language model training
[AAAI 2025] Assessing the Creativity of LLMs in Proposing Novel Solutions to Mathematical Problems
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
A collection of LLM papers, blogs, and projects, with a focus on OpenAI o1 🍓 and reasoning techniques.
Merging Generated and Retrieved Knowledge for Open-Domain QA (EMNLP 2023)
Official repository for ICML 2024 paper "On Prompt-Driven Safeguarding for Large Language Models"
The repository for ACL 2024 paper "TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models"