-
Sun Yat-sen University
- kongzhecn.github.io
Lists (6)
Sort Name ascending (A-Z)
Stars
[ICML2026] Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
Official repo for paper "Echo-Infinity: Learnable Evolving Memory for Real-Time Infinite Video Generation"
A Simple Baseline for Video World Models with Memory
Code for the paper HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement
[SIGGRAPH Asia 26 Conditionally Accept]PAct: Part-Decomposed Single-View Articulated Object Generation
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Infinite Worlds with Versatile Interactions
[ECCV 2026] WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG
Code for MIRA: Multiplayer Interactive World Models with Representation Autoencoders
EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis
DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation
AI PPT赛道终结者,史上最最最强 PPT Skill!!! 使用GPT生成豪华的图片格式PPT,然后转换为完全可编辑的PPTX文件。
Wan: Open and Advanced Large-Scale Video Generative Models
A toolkit for speaker diarization.
Official Implementation of LongLive-RAG: A general retrieval-augmented framework for long video generation.
JoyAI-Echo: Pushing the Frontier of Long Audio-Visual Generation
Official page of ImmerIris: A Large-Scale Dataset and Benchmark for Off-Axis and Unconstrained Iris Recognition in Immersive Applications.
"CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub: https://clianything.cc/
Implementation of Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
Codex skill for converting slide images, PDFs, and image-based PPTX files into editable PowerPoint decks.
Multimodal RL training framework for diffusion & omni models
Official Repo of "D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models"
A unified framework for easy reinforcement learning in Flow-Matching models
AI turns documents or topics into real, native PowerPoint decks—with native shapes, transitions and animations, data-backed charts and tables on demand, audio narration from speaker notes, and supp…
Interactive World Model papers organized by core research challenges.
[ICML 2026] World-R1: Reinforcing 3D Constraints for Text-to-Video Generation