Lists (4)
Sort Name ascending (A-Z)
Stars
A clean single-column LaTeX preprint and technical report template with elegant typography and modern front matter.
🔥 DanceOPD: On-Policy Generative Field Distillation
[ICML 2026] LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts
Repo for SeedVR2 (ICLR2026) & SeedVR (CVPR2025 Highlight)
Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
Infinite Interactive World Rollout on a Single Desktop GPU
Official implementation of "WorldKV: Efficient World Memory with World Retrieval and Compression"
Next Forcing: World Action Modeling with Multi-Chunk Prediction (MCP)
Infinite Worlds with Versatile Interactions
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Official Repository for Sakuga-42M Dataset
Tiny AutoEncoder for Hunyuan Video (and other video models)
Data-Forcing Distillation (DFD): restoring diversity and fidelity in few-step video generation — text-to-video (Wan2.1) & image-to-video (Cosmos), built on NVIDIA FastGen.
Official implementation of "MilliVid: Adaptive Latents for Long-Range Consistency in Video Generation"
Official Pytorch Code of the Paper "FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization"
Native Multimodal Models are World Learners
Learn everything about world model
Implementation of Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained GPUs.
DreamX-World: A General-Purpose Interactive World Model
Interactive World Model papers organized by core research challenges.
SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia
Skills for Real Engineers. Straight from my .agents directory.
Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
[CVPR 2026] WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories (WorldExpand of HY-World 2.0)