-
University of Maryland
- College Park, MD, US
- https://yu-fangxu.github.io/
Highlights
- Pro
Stars
slime is an LLM post-training framework for RL Scaling.
🚀 Ultra Recipe for Training Long-Horizon Search Agents - matching frontier AI's search capability with a 20B model + stateful harness
Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration.
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours
AI agents running research on single-GPU nanochat training automatically
Executable, measurable, and reproducible AI4AI toward recursive self-improvement. Home of OpenMLE and Frontis-MA1.
A curated list of autonomous improvement loops, research agents, and autoresearch-style systems inspired by Karpathy's autoresearch.
[COLM 2026] TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning
A user-friendly & efficient knowledge distillation framework for LLMs, supporting off-policy, on-policy (OPD), cross-tokenizer, multimodal, and on-policy self-distillation.
A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
A curated collection of papers and resources on On-Policy Distillation for Large Language Models.
A curated list of resources (surveys, papers, benchmarks, and opensource projects) on Rubrics
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Paper list of agent for science
[ICLR 2026] Official repo for "Spotlight on Token Perception for Multimodal Reinforcement Learning"
[Findings of ACL 2026] ArrowGEV: Grounding Events in Video via Learning the Arrow of Time