Stars
Official implementation of EVOKE: Endless Interactive World with Bounded State and Long-Horizon Supervision. A three-step, CFG-free interactive world model. SOTA on WBench.
[Tech Report] Context Scaling: Scaling Properties of Text Conditioning in Visual Generation
Implementation of CS336 Assignment1
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
Self Gradient Forcing (SGF) recovers the missing context-gradient path for self-generated causal memory through a bounded two-pass replay, enabling models trained with only a 5-second window to ext…
Manipulation Skill Framework, an open source GPU parallelized robotics simulator and benchmark
Official code for paper Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
Infinite Worlds with Versatile Interactions
The job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.
A Scientific Way - Solve 2000 LeetCode Problems in 8 Months
My 2-Year College Notes Hub — Personal Web · 1000+ Blogs
Code for MIRA: Multiplayer Interactive World Models with Representation Autoencoders
Self-supervised learning for spatial perception
基于多智能体LLM的中文金融交易框架 - TradingAgents中文增强版
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale
Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.
JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System
[Official Code] PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory
DreamX-World: A General-Purpose Interactive World Model
Official code of Memento: Reconstruct to Remember for Consistent Long Video Generation
Code Release for "OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data"
🔥Multi-View Subject-Consistent Video Generation (SIGGRAPH 2026)
Solve puzzles. Improve your pytorch.
Implementation of Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.