-
University of Chinese Academy of Sciences
- Beijing, China
- zebinx.github.io
- https://scholar.google.com/citations?user=Fs9_PskAAAAJ
Stars
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
iMac: Translating Actions into Motion and Contact Images for Embodied World Models
[ICLR 2026 Oral] Latent Particle World Models official repository
[ICRA 2025] Towards Safe End-to-end Autonomous Driving via Online Map Uncertainty
Wan: Open and Advanced Large-Scale Video Generative Models
Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
SE-Agent is a self-evolution framework for LLM Code agents. It enables trajectory-level evolution to exchange information across reasoning paths via Revision, Recombination, and Refinement, expandi…
[CVPR 2026 Oral] Learning to Drive via Real-World Simulation at Scale
Unofficial implementation of Titans, SOTA memory for transformers, in Pytorch
The official repo of "Mimir: Hierarchical Goal-Driven Diffusion with Uncertainty Propagation for End-to-End Autonomous Driving"
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
[CVPR 2026] G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
Official implementation for DSRL, Steering Your Diffusion Policy with Latent Space Reinforcement Learning (CoRL 2025)
Official implementation of the paper "HUGSIM: A Real-Time, Photo-Realistic and Closed-Loop Simulator for Autonomous Driving"
MemGen: Weaving Generative Latent Memory for Self-Evolving Agents
A curated list of state-of-the-art research in embodied AI, focusing on vision-language-action (VLA) models, vision-language navigation (VLN), and related multimodal learning approaches.
Implementation of [CVPR 2025] "DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation"
[ICLR 2026] ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
[CVPR2022] Remember Intentions: Retrospective-Memory-based Trajectory Prediction
HE-Drive: Human-Like End-to-End Driving with Vision Language Models
MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flexible speaker control, and multilingual support, while enablin…
A curated list of awesome papers on Embodied AI and related research/industry-driven resources.
Lumina Robotics Talent Call | Lumina社区具身智能招贤榜 | A list for Embodied AI / Robotics Jobs (PhD, RA, intern, etc