Skip to content
View yunx-z's full-sized avatar

Highlights

  • Pro

Block or report yunx-z

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Welcome to KernelBench-Verified. This repository provides a robust, realistic evaluation framework for assessing custom CUDA kernels generated by Large Language Models (LLMs).

Python 11 Updated Jul 17, 2026

[ICML 2026] AdaMEM: Test-Time Adaptive Memory for Language Agents

Python 7 1 Updated Jun 22, 2026

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 194,961 109,378 Updated Jun 26, 2026

A single-file implementation of KV cache paged attention

Python 2 Updated Mar 18, 2026

mini-swe-agent-plus: a tiny (~100 LOC) GitHub issue fixer—now with a robust multi-line text edit tool.

Python 25 5 Updated Jan 20, 2026
Python 10 11 Updated Nov 14, 2025

A Practitioner's Guide to M(eow)ti Turn Agentic ReinfOrcement learning

Python 83 12 Updated Jan 16, 2026

[NeurIPS 2025 D&B Track] MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research

Python 33 4 Updated May 8, 2026

The official implementation of "ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering"

Python 72 5 Updated Jun 21, 2025

About The official GitHub page for ''Unlocking General Long Chain-of-Thought Reasoning Capabilities of Large Language Models via Representation Engineering'' Resources

Python 11 3 Updated Jun 20, 2025

An AI Hedge Fund Team

Python 62,544 11,004 Updated Jul 31, 2026

Pax Renaissance for Board Game Arena

PHP 2 1 Updated Jun 8, 2025

MLGym A New Framework and Benchmark for Advancing AI Research Agents

Python 613 59 Updated Aug 10, 2025

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 31,020 7,559 Updated Aug 1, 2026

This repository contains the 1st place source code of Track II: Backdoor Trigger Recovery for Models in The Competition for LLM and Agent Safety 2024 at NeurIPS 2024.

Python 3 1 Updated Nov 3, 2024

A curated list of papers on LLMs and agents for scientific research and development

97 6 Updated Dec 11, 2024

Official code for paper: Chain of Ideas: Revolutionizing Research via Novel Idea Development with LLM Agents

Python 509 30 Updated Jan 15, 2025

DSIR large-scale data selection framework for language model training

Python 275 19 Updated Apr 7, 2024

[AAAI 2025] Assessing the Creativity of LLMs in Proposing Novel Solutions to Mathematical Problems

Jupyter Notebook 13 4 Updated May 5, 2025

MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering

Python 1,659 258 Updated Apr 24, 2026

OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Python 1,849 132 Updated Jan 17, 2025

A collection of LLM papers, blogs, and projects, with a focus on OpenAI o1 🍓 and reasoning techniques.

6,895 369 Updated Dec 17, 2025

Merging Generated and Retrieved Knowledge for Open-Domain QA (EMNLP 2023)

Python 21 1 Updated Oct 8, 2023

Official repository for ICML 2024 paper "On Prompt-Driven Safeguarding for Large Language Models"

Python 108 9 Updated May 20, 2025
Python 12 1 Updated Jan 25, 2024

The repository for ACL 2024 paper "TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models"

Python 36 2 Updated Jun 29, 2024
Python 25 Updated Jun 10, 2025
Python 2 Updated Jul 12, 2022
Next