Skip to content
View tianjianl's full-sized avatar

Block or report tianjianl

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

[ICLR 2026] Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs

Python 119 10 Updated Feb 2, 2026

An LLM post-training framework with vLLM for RL Scaling

Python 411 74 Updated Aug 3, 2026

Helpful tools and examples for working with flex-attention

Python 1,224 77 Updated Aug 11, 2026

Research artifacts from Recursive's automated AI research system

Python 203 19 Updated Jun 11, 2026

Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]

Python 37 7 Updated May 31, 2026

Programmable chat templates for LLM training and inference.

Python 144 31 Updated Aug 11, 2026

Framework for evaluating and improving agents

Python 4,120 1,525 Updated Aug 11, 2026

KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels

Python 46 2 Updated Jun 1, 2026
Python 370 46 Updated Jun 9, 2026

AIRS-Bench: an AI Research Science benchmark for quantifying the end-to-end AI research abilities of LLM agents

Python 108 9 Updated May 5, 2026

Can Language Models Rebuild Programs From Scratch?

Python 885 62 Updated Jul 26, 2026

Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.

Python 364 86 Updated Aug 11, 2026

Ship correct and fast LLM kernels to PyTorch

Python 153 18 Updated Jan 14, 2026

AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solution.

Python 532 147 Updated Aug 11, 2026
Python 10 1 Updated Apr 27, 2026

A lightweight, AI-native training framework for large language models. Designed for fast iteration, reproducible experiments, and modular configuration across SFT, RLVR, and evaluation workflows.

Python 583 43 Updated May 18, 2026

Harness for running and evaluating AI agents against RL environments

Python 232 53 Updated Aug 11, 2026

Training tiny models to prove hard theorems

Python 81 15 Updated Mar 5, 2026

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

Python 1,959 352 Updated Aug 11, 2026

Benchmarking Language Agents Under Controllable and Extreme Context Growth

Python 52 9 Updated Apr 29, 2026

[KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML 2026]

Python 202 33 Updated Mar 29, 2026

The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in language modeling.

Jupyter Notebook 145 15 Updated May 6, 2026

Dated Data: Tracing Knowledge Cutoffs in Large Language Models

Python 7 Updated Aug 11, 2025

General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.

Python 5,432 881 Updated Aug 9, 2026

Super basic implementation (gist-like) of RLMs with REPL environments.

Python 842 139 Updated Jan 7, 2026

Material for gpu-mode lectures

Jupyter Notebook 6,424 642 Updated Jun 15, 2026

A Lightweight LLM Post-Training Library

Python 2,398 330 Updated Aug 11, 2026
Next