Highlights
- Pro
Lists (1)
Sort Name ascending (A-Z)
Stars
Companion code for the global workspace interpretability paper
Three reward hacking environments: code, medical chat, biography generation. This repo contains code for the paper "Designing Effective Monitor-Based Interventions for Mitigating Reward Hacking Dur…
Official Inspect Implementation for "ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases"
🚀 Ultra Recipe for Training Long-Horizon Search Agents - matching frontier AI's search capability with a 20B model + stateful harness
Retrieval is CheapShow Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation
ICML 2026 · Plug-and-play long-term memory for LLM agents
Fully open data curation for reasoning models
Fast, accurate & comprehensive text measurement & layout
DFlash: Block Diffusion for Flash Speculative Decoding
Implementation for FP8/INT8 Rollout for RL training without performence drop.
BioDSA: Framework for Vibe Prototyping of AI Agents for Biomedicine
This repository is an official code base for our paper, Randomly Removing 50% of Dimensions in Text Embeddings has Minimal Impact on Retrieval and Classification Tasks (EMNLP main, 2025).
This repository contains the toolkit for replicating results from our technical report.
Generative AI for designing easily synthesizable small molecule drugs
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)
Repo for "Adaptation of Agentic AI"
A Survey of Reinforcement Learning for Large Reasoning Models
The code for paper "EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning"
ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program Evolution 🧬
On the Theoretical Limitations of Embedding-Based Retrieval
JARVIS, a system to connect LLMs with ML community. Paper: https://arxiv.org/pdf/2303.17580.pdf
DSPy: The framework for programming—not prompting—language models
TextGrad: Automatic ''Differentiation'' via Text -- using large language models to backpropagate textual gradients. Published in Nature.