Lists (12)
Sort Name ascending (A-Z)
Stars
A corruption robustness diagnostic testbed for vision-language models
将博导十年科研经验炼化为可直接调用的 AI 技能。从 Idea 构思到论文投稿,你的 AI 科研副导师。
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
QuantClaw is a plug-and-play task-type routing quantization plugin for OpenClaw.
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
[AAAI 2026] Beyond Cosine Similarity: Magnitude-Aware CLIP for No-Reference Image Quality Assessment
一个基于nano banana pro🍌的原生AI PPT生成应用,迈向"Vibe PPT"; 支持上传任意模板图片,上传任意素材&智能解析,一句话/大纲/页面描述自动生成PPT,口头修改指定区域、一键导出可编辑ppt - An AI-native slides generator based on nano banana pro🍌
A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing
[NeurIPS 2025] Official PyTorch implementation of paper "BADiff: Bandwidth Adaptive Diffusion Model"
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
**Deep Video Discovery (DVD)** is a deep-research style question answering agent designed for understanding extra-long videos.
Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)
This repository contains low-bit quantization papers from 2020 to 2026 on top conference.
[NeurIPS 2025 Spotlight] VisualQuality-R1 is the first open-sourced NR-IQA model can accurately describe and rate the image quality.
Official implementation of GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
Q-Insight is open-sourced at https://github.com/bytedance/Q-Insight. This repository will not receive further updates.
Beyond Accuracy: What Matters in Designing Well-Behaved Models?
Janus-Series: Unified Multimodal Understanding and Generation Models
[Paper List‘25] Paper List of Visual Data Coding for Machines, including Image/Video Coding for Machines, Feature Compression, Point Cloud Compression for Machines and Image/Video Coding for Machin…
Model Compression Toolbox for Large Language Models and Diffusion Models
Official codes for "Q-Ground: Image Quality Grounding with Large Multi-modality Models", ACM MM2024 (Oral)
Official repo for `LMM-PCQA: Assisting Point Cloud Quality Assessment with LMM', ACM MM2024 Oral
h4nwei / Compare2Score
Forked from Q-Future/Compare2Score[NeurIPS'24] Compare2Score