-
HuaZhong University of Science and Technology
- WuHan
Stars
MAGI-1: Autoregressive Video Generation at Scale
[CVPR 2025] GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding
The first decoder-only multimodal state space model
Pathology Foundation Model - Nature Medicine
finetuning SAM with non-promptable decoder on medical images
From a video, automatically create an Instance Segmentation dataset using Detectors like YoloX and Segment Anything
[MedIA'25] UN-SAM: Domain-Adaptive Self-Prompt Segmentation for Universal Nuclei Images
[AAAI 2025] Linear-complexity Visual Sequence Learning with Gated Linear Attention
[ICLR 2025] This is the official repository of our paper "MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine“
Segment Anything in Medical Images
The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
[KBS'25] NuSegDG: Integration of Heterogeneous Space and Gaussian Kernel for Domain-Generalized Nuclei Segmentation
[CVPR 2025 Highlight] Truncated Diffusion Model for Real-Time End-to-End Autonomous Driving
[ECCV 2024] Code for "Unleashing the Power of Prompt-driven Nucleus Instance Segmentation"
Bridging Large Vision-Language Models and End-to-End Autonomous Driving
Swin-LiteMedSAM: A Lightweight Box-Based Segment Anything Model for Large-Scale Medical Image Datasets
Accepted in CVPR 2023
PySlowFast: video understanding codebase from FAIR for reproducing state-of-the-art video models.
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use th…
[arXiv '24] Efficient Cell Nuclei Instance Segmentation with Large Convolution Kernels
Code for CVPR25 paper "VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos"
Video Highlight generation using short time analysis and keyframe algorithm
[ECCV 2024🔥] Official implementation of the paper "ST-LLM: Large Language Models Are Effective Temporal Learners"
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
mllm-npu: training multimodal large language models on Ascend NPUs
Official Repository of paper VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding