Skip to content
View Yu-Fangxu's full-sized avatar

Highlights

  • Pro

Organizations

@tianyi-lab

Block or report Yu-Fangxu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Weak-to-Strong On-Policy Distillation

11 Updated Jul 30, 2026

Open Frontier Intelligence

7,677 528 Updated Jul 28, 2026

[COLM 2026] TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning

Python 7 Updated Jul 13, 2026
Jupyter Notebook 17 1 Updated May 31, 2026
Python 2 Updated Jul 10, 2026

A user-friendly & efficient knowledge distillation framework for LLMs, supporting off-policy, on-policy (OPD), cross-tokenizer, multimodal, and on-policy self-distillation.

Python 231 17 Updated Jul 30, 2026

A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models

591 20 Updated Jul 31, 2026

Awesome List for On-Policy Distillation

781 20 Updated Jul 31, 2026

A curated collection of papers and resources on On-Policy Distillation for Large Language Models.

Python 485 10 Updated Jul 31, 2026

A curated list of resources (surveys, papers, benchmarks, and opensource projects) on Rubrics

103 4 Updated Jul 29, 2026

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Python 491 94 Updated May 18, 2026

SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles

Python 4,404 379 Updated Jul 31, 2026

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Python 869 59 Updated Jun 29, 2026

Paper list of agent for science

286 25 Updated Jun 27, 2026

[ICLR 2026] Official repo for "Spotlight on Token Perception for Multimodal Reinforcement Learning"

Python 71 7 Updated Apr 3, 2026

[Findings of ACL 2026] ArrowGEV: Grounding Events in Video via Learning the Arrow of Time

Python 4 Updated Apr 19, 2026

MiroEval: A benchmark and evaluation framework for deep research agents — 100 tasks (70 text, 30 multimodal) assessed across synthesis quality, factuality, and research process. 13 systems evaluated.

Python 46 7 Updated Jul 6, 2026

Awesome Unified Multimodal Models

1,306 46 Updated Mar 24, 2026

A unified multimodal model toolkit

Python 452 62 Updated Jul 29, 2026

[ICML 2026] XSkill: Continual Learning from Experience and Skills in Multimodal Agents

Python 240 32 Updated May 13, 2026

Official repository for the paper "Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation"

Python 277 15 Updated May 28, 2026

Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation

Python 256 21 Updated Apr 13, 2026

SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Python 918 71 Updated May 17, 2026

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 384,709 80,856 Updated Jul 31, 2026

[ICLR 2025] "GraphRouter: A Graph-based Router for LLM Selections", Tao Feng, Yanzhen Shen, Jiaxuan You

Python 75 7 Updated Dec 30, 2025

Qwen3.6 is the large language model series developed by Qwen team, Alibaba Group.

3,747 261 Updated Jun 3, 2026

ICLR 2026 (Oral) | EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

Python 59 4 Updated Feb 12, 2026
3 Updated Jan 30, 2026
Python 4,588 500 Updated Apr 22, 2026

OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis

Python 1,117 102 Updated Jun 10, 2026
Next