Skip to content
View SiLangWHL's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report SiLangWHL

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.

Python 1,956 220 Updated Jul 25, 2026

A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Python 1,027 35 Updated Jun 16, 2026

DeepTumorVQA benchmark for VLMs and Agents (10k testing samples)

Python 41 1 Updated May 19, 2026

[ICLR 26] TempFlow-GRPO (Temporal Flow GRPO), a principled GRPO framework that captures and exploits the temporal structure inherent in flow-based generation.

Python 481 25 Updated Nov 24, 2025

[EMNLP 2025 Industry] Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning

Python 36 Updated Oct 22, 2025
Python 572 33 Updated Nov 26, 2024

The code for "TokenPacker: Efficient Visual Projector for Multimodal LLM", IJCV2025

Python 279 9 Updated May 26, 2025

[ICLR '25] Official Pytorch implementation of "Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations"

Python 105 9 Updated Nov 30, 2025

A paper list of some recent works about Token Compress for Vit and VLM

946 46 Updated Aug 10, 2026

[ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video

Python 1,849 98 Updated Nov 28, 2025

[CVPR 2022] Official Pytorch code for OW-DETR: Open-world Detection Transformer

Python 263 46 Updated Apr 4, 2023

Official implementation of "FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on"

Python 627 59 Updated Feb 8, 2025

[ICLR 2025] CatVTON is a simple and efficient virtual try-on diffusion model with 1) Lightweight Network (899.06M parameters totally), 2) Parameter-Efficient Training (49.57M parameters trainable) …

Python 1,819 236 Updated Dec 16, 2025

Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Jupyter Notebook 4,293 363 Updated Nov 27, 2025

The ultimate training toolkit for finetuning diffusion models

Python 11,679 1,478 Updated Aug 11, 2026

🌟Change the world, it will become a better place. | 以科研和竞赛为导向的最好的YOLO实践框架!

Python 247 38 Updated Oct 1, 2024

Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.

Python 12,951 1,705 Updated Jul 10, 2026

Model interpretability and understanding for PyTorch

Python 5,685 561 Updated Aug 11, 2026

[CVPR 2024] Official RT-DETR (RTDETR paddle pytorch), Real-Time DEtection TRansformer, DETRs Beat YOLOs on Real-time Object Detection. 🔥 🔥 🔥

Python 5,446 647 Updated Aug 11, 2026

A Large-Scale Dataset for Spinal Vertebrae Segmentation in Computed Tomography

Python 201 39 Updated May 25, 2025

Developing Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography

Python 116 11 Updated Oct 15, 2024

Developing Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography

Python 411 56 Updated Jul 18, 2025

[npj Digital Medicine] The official repository for "Large-Vocabulary Segmentation for Medical Images with Text Prompts"

Python 309 23 Updated Dec 29, 2025

MICCAI 2024: Tri-Plane Mamba: Efficiently Adapting Segment Anything Model for 3D Medical Images

Python 28 1 Updated Apr 3, 2025

[ECCV 2024] Scene as Gaussians for Vision-Based 3D Semantic Occupancy Prediction

Python 680 62 Updated Mar 18, 2025

The implementation of "In-Place Scene Labelling and Understanding with Implicit Scene Representation" [ICCV 2021].

Python 460 58 Updated Jun 16, 2023

[MICCAI 2024] Learning 3D Gaussians for Extremely Sparse-View Cone-Beam CT Reconstruction

Python 65 5 Updated Mar 9, 2025

A Benchmark Dataset: Synthesized DeepLesion for CT Metal Artifact Reduction

78 9 Updated Dec 7, 2025

[ICLR2025] Halton Scheduler for Masked Generative Image Transformer

Python 288 33 Updated Oct 28, 2025
Next