-
Tsinghua University
- Beijing, China
Stars
MAGI-2-preview: Scaling Video Generation Models Efficiently
FlashKDA: high-performance Kimi Delta Attention kernels
Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.
Infinite Worlds with Versatile Interactions
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Self-supervised learning for spatial perception
GPT Image 2 prompt gallery, image prompt library, agentic skill, and CLI for OpenAI image generation/editing
[ICML 2026] World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
WorldEngine: Towards the Era of Post-Training for Autonomous Driving
Official Code Release of SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
Information collection for the Happy Horse AI video generator model. Official demo and updates at happyhorses.io.
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
A library for efficient similarity search and clustering of dense vectors.
Masked Depth Modeling for Spatial Perception
Official PyTorch Implementation of "Latent Denoising Makes Good Visual Tokenizers"
Sharp Monocular View Synthesis in Less Than a Second
Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
Native and Compact Structured Latents for 3D Generation
[CVPR 2026] Official Implementation of Particulate: Feed-Forward 3D Object Articulation
[CVPR 2026] "E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training" official implementation.
A Cross-Platform Backend for High-Performance Sparse Convolutions
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
cuTile is a programming model for writing parallel kernels for NVIDIA GPUs