Skip to content
View Yu-Fangxu's full-sized avatar

Highlights

  • Pro

Organizations

@tianyi-lab

Block or report Yu-Fangxu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

slime is an LLM post-training framework for RL Scaling.

Python 8,049 1,149 Updated Aug 14, 2026
Python 142 8 Updated Mar 31, 2026

🚀 Ultra Recipe for Training Long-Horizon Search Agents - matching frontier AI's search capability with a 20B model + stateful harness

Python 982 153 Updated Jun 15, 2026

Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration.

Python 3,588 431 Updated Aug 16, 2026
Python 130 24 Updated Aug 1, 2026

MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering

Python 1,691 258 Updated Apr 24, 2026

KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)

Jupyter Notebook 1,199 189 Updated Mar 24, 2026

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

Python 515 58 Updated Aug 13, 2026
Python 5 Updated Aug 14, 2026

The open source coding agent.

TypeScript 197,922 25,488 Updated Aug 16, 2026

AI agents running research on single-GPU nanochat training automatically

Python 93,923 13,313 Updated Mar 26, 2026

Executable, measurable, and reproducible AI4AI toward recursive self-improvement. Home of OpenMLE and Frontis-MA1.

Python 482 45 Updated Aug 12, 2026

A curated list of autonomous improvement loops, research agents, and autoresearch-style systems inspired by Karpathy's autoresearch.

2,439 183 Updated Aug 10, 2026

Weak-to-Strong On-Policy Distillation

28 2 Updated Aug 9, 2026

Open Frontier Intelligence

8,469 660 Updated Aug 6, 2026

[COLM 2026] TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning

Python 7 Updated Jul 13, 2026
Jupyter Notebook 20 2 Updated May 31, 2026
Python 2 Updated Jul 10, 2026

A user-friendly & efficient knowledge distillation framework for LLMs, supporting off-policy, on-policy (OPD), cross-tokenizer, multimodal, and on-policy self-distillation.

Python 236 17 Updated Aug 16, 2026

A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models

671 21 Updated Aug 15, 2026

Awesome List for On-Policy Distillation

828 18 Updated Jul 31, 2026

A curated collection of papers and resources on On-Policy Distillation for Large Language Models.

Python 512 10 Updated Aug 12, 2026

A curated list of resources (surveys, papers, benchmarks, and opensource projects) on Rubrics

107 4 Updated Aug 8, 2026

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Python 500 95 Updated May 18, 2026

SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles

Python 4,838 411 Updated Aug 16, 2026

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Python 930 66 Updated Jun 29, 2026

Paper list of agent for science

288 26 Updated Jun 27, 2026

[ICLR 2026] Official repo for "Spotlight on Token Perception for Multimodal Reinforcement Learning"

Python 81 7 Updated Apr 3, 2026

[Findings of ACL 2026] ArrowGEV: Grounding Events in Video via Learning the Arrow of Time

Python 3 Updated Apr 19, 2026
Next