-
Shanghai Jiao Tong University
- 800 Dongchuan Road, Shanghai, China
Highlights
- Pro
Starred repositories
A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
Mechanistic Interpretability toolkit for Vision-Language-Action models
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data
🍑 relsim: Relational Visual Similarity | pip install relsim 🌍 (CVPR 2026)
启智平台任务管理 CLI:资源查询、任务提交、日志查看和 MCP/agent workflow
Official Repo of The Great March Project. https://www.rhos.ai/research/gm-100
[AAAI 2026] The Official Implementation for "Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation"
[CSUR] A Survey on Video Diffusion Models
[ICCV 2025] Official implementation of "Holistic Unlearning Benchmark: A Multi-Faceted Evaluation for Text-to-Image Diffusion Model Unlearning"
A comprehensive reading list for Emotion Recognition in Conversations
awesome papers in LLM interpretability
[TMLR 2025🔥] A survey for the autoregressive models in vision.
collection of diffusion model papers categorized by their subareas
😎 up-to-date & curated list of awesome 3D Visual Grounding papers, methods & resources.
EO: Open-source Unified Embodied Foundation Model Series
Microsoft BASIC for 6502 Microprocessor - Version 1.1
A resource repository for machine unlearning in large language models
A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository aggregates surveys, blog posts, and research papers that explor…
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.
Code for "Scaling Language-Free Visual Representation Learning" paper (Web-SSL).
A curated list of state-of-the-art research in embodied AI, focusing on vision-language-action (VLA) models, vision-language navigation (VLN), and related multimodal learning approaches.
Code for Scaling Language-Free Visual Representation Learning (WebSSL)
An aggregation of human motion understanding research.