-
Beijing University of Posts and Telecommunications
- Beijing
-
03:06
(UTC +08:00) - https://zzzyzh.github.io/
- https://scholar.google.com/citations?user=CMzNexYAAAAJ&hl=zh-CN
Highlights
- Pro
Lists (9)
Sort Name ascending (A-Z)
Stars
Code for kai0, including training, inference and data collection.
Official style files for papers submitted to venues of the Association for Computational Linguistics
Official Repo of "Flow-OPD: On-Policy Distillation for Flow Matching Models"
[ACM MM 2026] On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-Image Generation
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation, The Fourteenth International Conference on Learning Representations (ICLR) 2026, Accepted
OPRD: On-Policy Representation Distillation (https://arxiv.org/abs/2606.06021)
openpi-RLT is an openpi-based real-robot RL system with RL-token-guided action refinement.
[RSS 2026] Code for RISE: Self-Improving Robot Policy with Compositional World Model
Codex-native Academic Research Skills suite for human-in-the-loop academic research workflows
Academic Research Skills for Claude Code: research → write → review → revise → finalize
[ECCV 2026] EgoSim: Egocentric World Simulator for Embodiment Interaction Generation
NEO Series: Native Vision-Language Models from First Principles
The implementation of Teachability-Aware On-Policy Distillation (TA-OPD), the method introduced in Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation.
Everything about the SmolLM and SmolVLM family of models
On Policy Distillation Build on top of Verl
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Code, data and weights for the paper **What drives success in physical planning with Joint-Embedding Predictive World Models?**
Official repository for the paper "Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation"
Post-training with Tinker
HY-Embodied: Embodied Foundation Models for Real-World Agents
Official implementation of "OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning"
Source code for the Refined Policy Distillation paper.
Vision-OPD is a regional-to-global on-policy self-distillation framework that transfers a model's own privileged crop-conditioned perception to its full-image policy, enabling fine-grained visual u…
A high-throughput and memory-efficient inference and serving engine for LLMs