-
Institute for Intelligent Computing, Alibaba Group
- Hangzhou, China
- https://scholar.google.com/citations?user=GHOQKCwAAAAJ&hl=zh-CN
Stars
Infinite Interactive World Rollout on a Single Desktop GPU
[SIGGRAPH(TOG)'2026] ATGS: Anchored Temporal Gaussian Splatting for Long Volumetric Video Representation
Reimplementation of LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
The project involves producing a basketball game highlights movie utilizing yolov5 for detecting the basket and resnet50 for identifying scoring actions
[KDD 2026 Datasets & Benchmarks Track] SVHighlights: a benchmark for highlight detection in extremely long sports videos
SPEAR: A Simulator for Photorealistic Embodied AI Research
Towards a Generative 3D World Engine for Embodied Intelligence
[ECCV 2026] OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
A feed-forward 3D foundation model for reconstructing scenes from streaming data
[TMM 2023] StrongSORT: Make DeepSORT Great Again
Official code for MAMMA: Markerless Accurate Multi-person Motion Acquisition.
Helios: Real Real-Time Long Video Generation Model
ViGeo: Towards Consistent Video Geometry Estimation
Turn paper/text/topic into editable research figures, technical route diagrams, and presentation slides.
LPM 1.0: Video-based Character Performance Model
Open-source Windows and Office activator featuring HWID, Ohook, TSforge, and Online KMS activation methods, along with advanced troubleshooting.
OmX - Oh My codeX: Your codex is not alone. Add hooks, agent teams, HUDs, and so much more.
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
LHM++: An Efficient Large Human Reconstruction Model for Pose-free Images to 3D
Official code for "LagerNVS Latent Geometry for Fully Neural Real-time Novel View Synthesis" (CVPR 2026)
THEORY OF SPACE: a benchmark for evaluating whether foundation models can actively explore under partial observability efficiently to build, update, and exploit globally consistent spatial beliefs.
slime is an LLM post-training framework for RL Scaling.
Agent Laboratory is an end-to-end autonomous research workflow meant to assist you as the human researcher toward implementing your research ideas
[CVPR 2026] Towards Real-Time Diffusion-Based Streaming Video Super-Resolution — An efficient one-step diffusion framework for streaming VSR with locality-constrained sparse attention and a tiny co…
Official Pytorch Implementation of SMIRK: 3D Facial Expressions through Analysis-by-Neural-Synthesis (CVPR 2024)