-
Alibaba Group
- Beijing, Chaoyang
- https://vasgaowei.github.io/
- https://www.xiaohongshu.com/user/profile/626d1ddd0000000021021dc3
Stars
[Official Repo] JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
[Tech Report] Context Scaling: Scaling Properties of Text Conditioning in Visual Generation
MAGI-2-preview: Scaling Video Generation Models Efficiently
[ICML 2026] Multimodal deep-research MLLM and benchmark. The first long-horizon multimodal deep-research MLLM, extending the number of reasoning turns to dozens and the number of search-engine inte…
[Tech Report] Democratizing the Training of Video World Models from Scratch. 🔥 🔥 🔥
Official repository for the ACM MM 2026 paper “Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning”
[ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++
Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
Vision-OPD is a regional-to-global on-policy self-distillation framework that transfers a model's own privileged crop-conditioned perception to its full-image policy, enabling fine-grained visual u…
VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.
NVIDIA FastGen: Fast Generation from Diffusion Models
[arXiv 2026] This is the official PyTorch implementation of "MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators".
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning
[ECCV2026] Official Implementation of "VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement"
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
AI-Powered Agentic Repository Intelligence Platform
Boogu-Image-0.1 is an Apache-2.0 open-source image generation and editing model family that delivers near-closed-source performance with an order of magnitude less data.
Sutskever 30 implementations inspired by https://papercode.vercel.app/ | For Agents, use https://github.com/pageman/Sutskever-Agent | Polyglot / Multi-Backed version at https://github.com/pageman/s…
On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators