Stars
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models (ReChannel)
Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning
Official Repo of "D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models"
Open source code for ICLR 2026 Paper: Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
[ECCV 2026] Official Code of "Distribution Matching Distillation Meets Reinforcement Learning"
🔥 Latest advances in Video Object Segmentation (VOS) – papers, datasets, and projects.
End2End Virtual Try-on with Visual Reference, CVPR2026
[ICLR 2026] Deforming Videos to Masks: Flow Matching for Referring Video Segmentation (FlowRVS)
[ICLR 2026]FZOO: Fast Zeroth-Order Optimizer for Fine‑Tuning Large Language Models towards Adam‑Scale Speed
[CVPR 2025] Low-Biased General Annotated Dataset Generation
[ICLR 2026] Self-Representation Alignment for Diffusion Transformers (SRA)
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners (EMNLP 2024 Oral)
An implementation of basic data structure and algorithm
A simple LaTeX package for Mathematical Contest in Modeling (MCM)
[CVPR'23] Probability-based Global Cross-modal Upsampling for Pansharpening