-
College of Computer Science, Sichuan University
- ChengDu, China
-
13:57
(UTC +08:00) - https://scholar.google.com/citations?user=gZBz65gAAAAJ
Stars
Official repository for CVPR 2025 paper: OpenSDI: Spotting Diffusion-Generated Images in the Open World
Code for paper "TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images"
PyTorch implementation of InstructDiffusion, a unifying and generic framework for aligning computer vision tasks with human instructions.
[ICML 2026] PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks
[CVPR 2026] UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection
Efficient Universal Perception Encoder: a single on-device vision encoder with versatile representations that match or exceed specialized experts across multiple task domains.
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
implementation of the paper Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
[CVPR 2026] The official PyTorch implementation of the "Vision Transformer Needs More Than Registers".
[NeurIPS 2025] MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
[AAAI 2025]Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning
This is the official repository for the paper "MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning"
Codes for Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
Official Repository for "Glyph: Scaling Context Windows via Visual-Text Compression"
[NeurIPS 2025] IEAP: Image Editing As Programs with Diffusion Models
Scan the Hallucination Citation of Academic papers. Convert second-hand citation to official version
[CVPR 26 Findings] Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios
[CVPR 2025 Highlight] Official code and models for Encoder-only Mask Transformer (EoMT).
[ICLR 2026] The official repository for paper "ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning"
Displays the China Computer Federation (CCF) recommended rank of international conferences and journals in the dblp, Google Scholar, Connected Papers and and Web of Science search results.