Skip to content
View GWxuan's full-sized avatar

Highlights

  • Pro

Block or report GWxuan

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

From Foundation to Application

Python 741 62 Updated Aug 10, 2026

iMac: Translating Actions into Motion and Contact Images for Embodied World Models

Python 33 Updated Jun 21, 2026

[Arxiv 2026] Toward the Whole Picture: Accumulative Fingerprint Mapping and Reconstruction for Small-Area Mobile Sensors

MATLAB 3 Updated Jun 17, 2026

The official application of Identity-Consistent Multi-Pose Generation of Contactless Fingerprints

Python 6 1 Updated May 6, 2026

GesVLA: Gesture-Aware Vision-Language-Action Model with Embedded Representations

Python 29 1 Updated May 22, 2026

[CVPR 2026] AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation

Python 76 4 Updated May 24, 2026

[ACL 2026 Poster] Code and Benchmark for "Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision"

Python 21 Updated Jun 3, 2026

RoboChallenge Inference example code

Python 151 9 Updated Jun 10, 2026

The official code of Yume

Python 683 45 Updated Jan 14, 2026

Official Hardware Codebase for the Paper "BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities"

Python 144 11 Updated Nov 18, 2025

[ICCV 2025] D^3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection

Python 17 Updated Jul 11, 2026

[ICLR 2026] Code of "MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation"

Python 319 25 Updated Jun 13, 2026

Codes, datasets, and synthetic dataset generator about the paper "LiCamPose: Combining Multi-View LiDAR and RGB Cameras for Robust Single-Snapshot 3D Human Pose Estimation"

Python 17 Updated Feb 28, 2026

Nav-R1: Reasoning and Navigation in Embodied Scenes

Python 131 4 Updated Oct 31, 2025

[CoRL 2025] GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation

Python 79 2 Updated Jun 21, 2026

[ICRA 2026] Official implementation of the paper: "StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling"

Python 576 47 Updated Nov 2, 2025
Python 59 2 Updated Aug 18, 2025

[ICCV 2025] IGL-Nav: Incremental 3D Gaussian Localization for Image-goal Navigation

68 Updated Aug 4, 2025

RoboBrain 2.5: Advanced version of RoboBrain. Depth in Sight, Time in Mind. πŸŽ‰πŸŽ‰πŸŽ‰

Python 1,126 113 Updated Feb 28, 2026
Python 257 8 Updated May 12, 2025

VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.

Python 3,853 330 Updated Mar 12, 2026

[CoRL 2025] Repository relating to "TrackVLA: Embodied Visual Tracking in the Wild"

Python 423 30 Updated Nov 25, 2025

Official implementation of "OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning"

Python 236 14 Updated May 30, 2025

[CoRL25] GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data

Python 389 16 Updated Dec 29, 2025

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Python 532 24 Updated Mar 2, 2026

[ICML 2025 Oral] Official repo of EmbodiedBench, a comprehensive benchmark designed to evaluate MLLMs as embodied agents.

Python 328 40 Updated May 30, 2026

[WACV'24] TD3D: Top-Down Beats Bottom-Up in 3D Instance Segmentation

Python 12 1 Updated Jun 7, 2025

[RSS'25] This repository is the implementation of "NaVILA: Legged Robot Vision-Language-Action Model for Navigation"

Python 689 71 Updated Aug 20, 2025

Code for OctoNav-Bench and OctoNav-R1

Python 77 2 Updated Apr 29, 2026

This is a PyTorch implementation of MCLN proposed by our paper "Multi-branch Collaborative Learning Network for 3D Visual Grounding"(ECCV2024)

Python 27 Updated Oct 10, 2024
Next