Skip to content
View h4nwei's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report h4nwei

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A corruption robustness diagnostic testbed for vision-language models

Python 2 Updated Jun 28, 2026

本人的科研经验

13,429 672 Updated Jun 6, 2026
Python 4 Updated Jul 16, 2026

将博导十年科研经验炼化为可直接调用的 AI 技能。从 Idea 构思到论文投稿,你的 AI 科研副导师。

Python 4,380 307 Updated Jul 16, 2026

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

Python 207 17 Updated May 16, 2026

QuantClaw is a plug-and-play task-type routing quantization plugin for OpenClaw.

TypeScript 115 1 Updated Apr 27, 2026

An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.

Python 1,943 220 Updated Jul 25, 2026
HTML 2 Updated Jun 18, 2026

[AAAI 2026] Beyond Cosine Similarity: Magnitude-Aware CLIP for No-Reference Image Quality Assessment

Python 13 1 Updated Nov 20, 2025

一个基于nano banana pro🍌的原生AI PPT生成应用,迈向"Vibe PPT"; 支持上传任意模板图片,上传任意素材&智能解析,一句话/大纲/页面描述自动生成PPT,口头修改指定区域、一键导出可编辑ppt - An AI-native slides generator based on nano banana pro🍌

TypeScript 15,323 1,766 Updated Jul 23, 2026

A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing

TypeScript 1,519 71 Updated Jul 3, 2026

[NeurIPS 2025] Official PyTorch implementation of paper "BADiff: Bandwidth Adaptive Diffusion Model"

14 1 Updated Oct 24, 2025

[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.

Python 735 78 Updated May 14, 2026

**Deep Video Discovery (DVD)** is a deep-research style question answering agent designed for understanding extra-long videos.

Python 404 20 Updated Nov 3, 2025

Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)

Python 727 30 Updated Sep 24, 2025

This repository contains low-bit quantization papers from 2020 to 2026 on top conference.

197 5 Updated Jun 25, 2026

[NeurIPS 2025 Spotlight] VisualQuality-R1 is the first open-sourced NR-IQA model can accurately describe and rate the image quality.

Python 201 8 Updated Oct 15, 2025

Official implementation of GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Python 252 19 Updated May 5, 2025

Q-Insight is open-sourced at https://github.com/bytedance/Q-Insight. This repository will not receive further updates.

142 3 Updated May 30, 2025

Beyond Accuracy: What Matters in Designing Well-Behaved Models?

Python 20 1 Updated Mar 30, 2026

Janus-Series: Unified Multimodal Understanding and Generation Models

Python 17,753 2,232 Updated Feb 1, 2025
8 Updated Jul 2, 2026

[Paper List‘25] Paper List of Visual Data Coding for Machines, including Image/Video Coding for Machines, Feature Compression, Point Cloud Compression for Machines and Image/Video Coding for Machin…

36 1 Updated Aug 17, 2025

Model Compression Toolbox for Large Language Models and Diffusion Models

Python 795 99 Updated Aug 14, 2025

Official codes for "Q-Ground: Image Quality Grounding with Large Multi-modality Models", ACM MM2024 (Oral)

Python 49 Updated Apr 21, 2026

Official repo for `LMM-PCQA: Assisting Point Cloud Quality Assessment with LMM', ACM MM2024 Oral

Python 17 1 Updated Nov 21, 2024

[NeurIPS'24] Compare2Score

Python 4 Updated Mar 24, 2026
Next