Lists (1)
Sort Name ascending (A-Z)
Stars
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
DeepTumorVQA benchmark for VLMs and Agents (10k testing samples)
[ICLR 26] TempFlow-GRPO (Temporal Flow GRPO), a principled GRPO framework that captures and exploits the temporal structure inherent in flow-based generation.
[EMNLP 2025 Industry] Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
The code for "TokenPacker: Efficient Visual Projector for Multimodal LLM", IJCV2025
[ICLR '25] Official Pytorch implementation of "Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations"
A paper list of some recent works about Token Compress for Vit and VLM
[ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
[CVPR 2022] Official Pytorch code for OW-DETR: Open-world Detection Transformer
Official implementation of "FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on"
[ICLR 2025] CatVTON is a simple and efficient virtual try-on diffusion model with 1) Lightweight Network (899.06M parameters totally), 2) Parameter-Efficient Training (49.57M parameters trainable) …
Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
The ultimate training toolkit for finetuning diffusion models
🌟Change the world, it will become a better place. | 以科研和竞赛为导向的最好的YOLO实践框架!
Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.
Model interpretability and understanding for PyTorch
[CVPR 2024] Official RT-DETR (RTDETR paddle pytorch), Real-Time DEtection TRansformer, DETRs Beat YOLOs on Real-time Object Detection. 🔥 🔥 🔥
A Large-Scale Dataset for Spinal Vertebrae Segmentation in Computed Tomography
Developing Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
Developing Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
[npj Digital Medicine] The official repository for "Large-Vocabulary Segmentation for Medical Images with Text Prompts"
MICCAI 2024: Tri-Plane Mamba: Efficiently Adapting Segment Anything Model for 3D Medical Images
[ECCV 2024] Scene as Gaussians for Vision-Based 3D Semantic Occupancy Prediction
The implementation of "In-Place Scene Labelling and Understanding with Implicit Scene Representation" [ICCV 2021].
[MICCAI 2024] Learning 3D Gaussians for Extremely Sparse-View Cone-Beam CT Reconstruction
A Benchmark Dataset: Synthesized DeepLesion for CT Metal Artifact Reduction
[ICLR2025] Halton Scheduler for Masked Generative Image Transformer