Stars
Sekai2: From World Exploration to Interactive World Modeling
On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
Vidu S1: A Real-Time Interactive Video Generation Model
ICLR'26: TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
[ICML 2026] The official implementation of paper "Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key to Unification"
Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
code for "Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion"
Code for RepWAM: World Action Modeling with Representation Visual-Action Tokenizers
Next Forcing: World Action Modeling with Multi-Chunk Prediction (MCP)
ARM: An AutoRegressive Large Multimodal Model with Discrete Representations
[CVPR 2026 Best Paper Finalist] Pixel Diffusion Transformers for Image Generation
[SIGGRAPH Asia 2026] DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
[ICML 2026] RoboTwin 2.0 Offical Code Repo
Benchmarking Knowledge Transfer in Lifelong Robot Learning
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
[ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Official Repo of "D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models"
slime is an LLM post-training framework for RL Scaling.
[Nips 2025] EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
Code accompanying Ego-Exo: Transferring Visual Representations from Third-person to First-person Videos (CVPR 2021)
HY-SOAR:Self-Correction for Optimal Alignment and Refinement in Diffusion Models
Official implementation of "OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning"
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
Skill package for ML/CV/NLP paper writing, curated and adapted from Prof. Peng Sida's open notes for Codex, Claude Code, and Gemini.