-
SJTU-SAI
- Shanghai
-
16:59
(UTC +08:00) - https://lixiny.github.io
- @lixinyang__
Highlights
- Pro
Lists (1)
Sort Name ascending (A-Z)
Stars
A multi-round co-design skill for publication-ready paper framework diagrams and method overview figures.
Official implementation of LaMP: Learning Vision-Language-Action Policy with 3D Scene Flow as Latent Motion Prior.
Official implementation of ChronoFlow-Policy: a diffusion-based visuomotor policy that jointly models past-current-future object-gripper interaction flows for robot manipulation.
[ICLR 2026] Codebase for paper "Geometry-aware 4D Video Generation for Robot Manipulation"
PaperBanana: Automating Academic Illustration For AI Scientists
The official implementation of InfiniteVGGT
Official Repo of The Great March Project. https://www.rhos.ai/research/gm-100
[CVPR 2026] InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields
🔥(CVPR 2025 Highlight) Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camera
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
The official repo for [NeurIPS'22] "ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation" and [TPAMI'23] "ViTPose++: Vision Transformer for Generic Body Pose Estimation"
ICCV 2025 | TesserAct: Learning 4D Embodied World Models
[ICRA 2026] VITRA: Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
Code for "EgoX: Egocentric Video Generation from a Single Exocentric Video"
The repository provides code for running inference with the SAM 3D Body Model (3DB), links for downloading the trained model checkpoints and datasets, and example notebooks that show how to use the…
Momentum Human Rig is an anatomically-inspired parametric full-body digital human model developed at Meta. It includes: A parametric body skeletal model; A realistic 3D mesh skinned to the skeleton…
[ICLR 2026] A simple state update rule to enhance length generalization for CUT3R
[ICLR 2026] Trace Anything: Representing Any Video in 4D via Trajectory Fields
Cosmos-Transfer2.5, built on top of Cosmos-Predict2.5, produces high-quality world simulations conditioned on multiple spatial control inputs.
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
[CoRL 2025] TWIST: Teleoperated Whole-Body Imitation System
[ICLR 2026] An unified model for 4D human-scene reconstruction
A lightweight suite of motion imitation methods for training controllers.
[arXiv 2025] VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
Unfied World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets