Skip to content
View wdrink's full-sized avatar

Block or report wdrink

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

This is a repo to track the latest autoregressive visual generation papers.

431 7 Updated Jun 25, 2025

Sekai2: From World Exploration to Interactive World Modeling (arXiv:2608.09449) — 128,892 real-world clips with camera trajectories and temporally grounded captions.

37 Updated Aug 14, 2026
Python 6,223 367 Updated Aug 15, 2026

On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

Python 545 43 Updated Aug 10, 2026

Open Frontier Intelligence

8,470 661 Updated Aug 6, 2026

Vidu S1: A Real-Time Interactive Video Generation Model

244 9 Updated Aug 6, 2026

ICLR'26: TTOM: Test-Time Optimization and Memorization for Compositional Video Generation

Python 7 Updated Mar 26, 2026

[ICML 2026] The official implementation of paper "Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key to Unification"

Python 50 Updated Jul 13, 2026

Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)

Python 3,478 281 Updated Sep 12, 2025

code for "Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion"

Python 1,285 71 Updated Jul 6, 2026

Code for RepWAM: World Action Modeling with Representation Visual-Action Tokenizers

64 1 Updated Aug 4, 2026

Next Forcing: Causal World Modeling with Multi-Chunk Prediction (MCP)

Python 115 3 Updated Aug 16, 2026

ARM: An AutoRegressive Large Multimodal Model with Discrete Representations

50 Updated Jun 10, 2026
Python 6 1 Updated Jun 9, 2026

[CVPR 2026 Best Paper Finalist] Pixel Diffusion Transformers for Image Generation

Python 924 73 Updated Jul 8, 2026

[SIGGRAPH Asia 2026] DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models

Python 158 2 Updated Jul 20, 2026

PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion

Python 1,033 61 Updated Jul 22, 2026

Flow Map OPD for AnyStep Video Diffusion

Python 411 11 Updated Aug 14, 2026

[ICML 2026] RoboTwin 2.0 Offical Code Repo

Python 2,727 457 Updated Aug 10, 2026

Benchmarking Knowledge Transfer in Lifelong Robot Learning

Jupyter Notebook 2,193 457 Updated Mar 15, 2025

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Python 1,963 308 Updated Jan 16, 2024

[ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process

Python 1,015 45 Updated Feb 10, 2026

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

Cuda 3,646 480 Updated Jan 17, 2026

Official Repo of "D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models"

Python 305 9 Updated May 22, 2026

slime is an LLM post-training framework for RL Scaling.

Python 8,051 1,149 Updated Aug 16, 2026

[Nips 2025] EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Python 144 2 Updated Jul 31, 2025

Code accompanying Ego-Exo: Transferring Visual Representations from Third-person to First-person Videos (CVPR 2021)

Python 40 5 Updated Jun 8, 2021

HY-SOAR:Self-Correction for Optimal Alignment and Refinement in Diffusion Models

Python 741 64 Updated Apr 21, 2026

Official implementation of "OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning"

Python 236 14 Updated May 30, 2025

Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models

Python 274 14 Updated Jul 23, 2026
Next