Skip to content
View Yu-Fangxu's full-sized avatar

Highlights

  • Pro

Organizations

@tianyi-lab

Block or report Yu-Fangxu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Weak-to-Strong On-Policy Distillation

6 Updated Jul 30, 2026

Open Frontier Intelligence

7,298 482 Updated Jul 28, 2026

[COLM 2026] TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning

Python 7 Updated Jul 13, 2026
Jupyter Notebook 17 1 Updated May 31, 2026
Python 2 Updated Jul 10, 2026

A user-friendly & efficient knowledge distillation framework for LLMs, supporting off-policy, on-policy (OPD), cross-tokenizer, multimodal, and on-policy self-distillation.

Python 230 17 Updated Jul 30, 2026

A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models

587 21 Updated Jul 29, 2026

Awesome List for On-Policy Distillation

774 20 Updated Jul 23, 2026

A curated collection of papers and resources on On-Policy Distillation for Large Language Models.

Python 480 10 Updated Jul 30, 2026

A curated list of resources (surveys, papers, benchmarks, and opensource projects) on Rubrics

103 4 Updated Jul 29, 2026

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Python 489 94 Updated May 18, 2026

SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles

Python 4,434 379 Updated Jul 29, 2026

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Python 861 58 Updated Jun 29, 2026

Paper list of agent for science

286 25 Updated Jun 27, 2026

[ICLR 2026] Official repo for "Spotlight on Token Perception for Multimodal Reinforcement Learning"

Python 70 7 Updated Apr 3, 2026

[Findings of ACL 2026] ArrowGEV: Grounding Events in Video via Learning the Arrow of Time

Python 4 Updated Apr 19, 2026

MiroEval: A benchmark and evaluation framework for deep research agents — 100 tasks (70 text, 30 multimodal) assessed across synthesis quality, factuality, and research process. 13 systems evaluated.

Python 46 7 Updated Jul 6, 2026

Awesome Unified Multimodal Models

1,306 46 Updated Mar 24, 2026

A unified multimodal model toolkit

Python 413 55 Updated Jul 29, 2026

[ICML 2026] XSkill: Continual Learning from Experience and Skills in Multimodal Agents

Python 239 31 Updated May 13, 2026

Official repository for the paper "Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation"

Python 276 15 Updated May 28, 2026

Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation

Python 256 21 Updated Apr 13, 2026

SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Python 916 71 Updated May 17, 2026

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 384,548 80,813 Updated Jul 30, 2026

[ICLR 2025] "GraphRouter: A Graph-based Router for LLM Selections", Tao Feng, Yanzhen Shen, Jiaxuan You

Python 75 7 Updated Dec 30, 2025

Qwen3.6 is the large language model series developed by Qwen team, Alibaba Group.

3,746 259 Updated Jun 3, 2026

ICLR 2026 (Oral) | EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

Python 59 4 Updated Feb 12, 2026
3 Updated Jan 30, 2026
Python 4,588 500 Updated Apr 22, 2026

OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis

Python 1,116 102 Updated Jun 10, 2026
Next