-
Johns Hopkins University
- Baltimore, MD
- tianjianl.github.io
- @tli104
Stars
[ICLR 2026] Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
An LLM post-training framework with vLLM for RL Scaling
Helpful tools and examples for working with flex-attention
Research artifacts from Recursive's automated AI research system
Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]
Programmable chat templates for LLM training and inference.
Framework for evaluating and improving agents
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
AIRS-Bench: an AI Research Science benchmark for quantifying the end-to-end AI research abilities of LLM agents
Can Language Models Rebuild Programs From Scratch?
Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.
Ship correct and fast LLM kernels to PyTorch
AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solution.
A lightweight, AI-native training framework for large language models. Designed for fast iteration, reproducible experiments, and modular configuration across SFT, RLVR, and evaluation workflows.
Harness for running and evaluating AI agents against RL environments
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
Benchmarking Language Agents Under Controllable and Extreme Context Growth
[KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML 2026]
The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in language modeling.
Dated Data: Tracing Knowledge Cutoffs in Large Language Models
General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.
Super basic implementation (gist-like) of RLMs with REPL environments.