- Hong Kong SAR
- @ComWjm
- https://jmwang.netlify.app/
Starred repositories
[RSS 2026] Learning Point Cloud Geometry as a Statistical Manifold: Theory and Practice
Official repository for the paper "ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation"
[ICML 2026] Official code release for paper "Temporal Score Rescaling for Temperature Sampling in Diffusion and Flow Models"
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
XRZero-G0: Pushing the Frontier of Dexterous Robotic Manipulation with Interfaces, Quality and Ratios
[ICML 2026] 🏂 World Guidance: World Modeling in Condition Space for Action Generation
[ICLR 2026] Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
Video-Action Models for Generalizable Robot Control Beyond VLAs
The agent that grows with you
Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps
A standlone demo from ManiDreams: Fast-FoundationStereo + SAM2 + Newton, zero-shot real-time simulation-based world model
Memory Sparse Attention - A scalable, end-to-end trainable latent-memory framework for 100M-token contexts.
Escaping the Big Data Paradigm in Self-Supervised Representation Learning
[ICRA 2026] VITRA: Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
[CVPR 2026 Highlight] ForeAct: Steering Your VLA with Efficient Visual Foresight Planning
[RSS 2026] Causal video-action world model for generalist robot control
[ICLR 2026] Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling
AgentFlow: In-the-Flow Agentic System Optimization
Running VLA at 30Hz frame rate and 480Hz trajectory frequency
A general-purpose robotic agent framework based on LLMs. The LLM can independently reason, plan, and execute actions to operate diverse robot types across various scenarios to complete unpredictabl…
Offical code release for DynoSAM: Dynamic Object Smoothing And Mapping. Accepted Transactions on Robotics (Visual SLAM SI). A visual SLAM framework and pipeline for Dynamic environements, estimatin…
1st place solution of 2025 BEHAVIOR Challenge
[CVPR 2026] Official codes of "Monet: Reasoning in Latent Visual Space Beyond Image and Language"
This repository contains code for the paper "Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training" by T. Bonnaire, R. Urfin, G. Biroli and M. Mézard.
(ICRA 2025) Inverse Mixed Strategy Games with Generative Trajectory Models
A highly robust and accurate LiDAR-only, LiDAR-inertial odometry
Official PyTorch Implementation of "Latent Denoising Makes Good Visual Tokenizers"
Visual Imitation Enables Contextual Humanoid Control. CoRL 2025, Best Student Paper Award.