Lists (2)
Sort Name ascending (A-Z)
Stars
极简单文件 Web 工具,用于运行/关闭/管理多个 rtunnel SSH 隧道。零依赖、跨设备同步设计。
MambaRaw: Selective State Space Modeling for Efficient 4K RAW Image Reconstruction (ECCV 2026)
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
[CVPR 2026] Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
[ICML2026] Imagination Helps Visual Reasoning, But Not Yet in Latent Space
An Open-Ended Benchmark and Formal Framework for Adjuvant Research with MLLMs
Towards Efficient Multimodal Large Language Models: A Survey on Token Compression
Practical Continual Forgetting for Pre-trained Vision Models (CVPR 2024; T-PAMI 2026)
Code and data for VTCBench, a VLM benchmark for long-context understanding capabilities under vision-text compression paradigm.
MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
[ICCV 2025] Official Implementation of Federated Continual Instruction Tuning
A Comprehensive Survey on Continual Learning in Generative Models.
[ACL'25 Main] Official Implementation of HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model
[ACL'25 Findings] ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing
[ICLR 2025 Oral] NeuralPlane: Structured 3D Reconstruction in Planar Primitives with Neural Fields
[ICLR 2025] Official Implementation of Local-Prompt: Extensible Local Prompts for Few-Shot Out-of-Distribution Detection
[NeurIPS'25 Spotlight🔥] Official Implementation of RobustMerge: Parameter-Efficient Model Merging for MLLMs with Direction Robustness
A collection of token reduction (token pruning, merging, clustering, etc.) techniques for ML/AI
(ICCV 2025) Enhance CLIP and MLLM's fine-grained visual representations with generative models.
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
[EMNLP'25 main] Official Implementation of ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
The codebase for paper "PPT: Token Pruning and Pooling for Efficient Vision Transformer"
[CVPR'25] Official Implementation of MambaIC: State Space Models for High-Performance Learned Image Compression