-
University of Macau
- Macau, China
-
14:24
(UTC +08:00) - http://www.zdzheng.xyz
- https://orcid.org/0000-0002-2434-9050
- https://scholar.google.com/citations?user=XT17oUEAAAAJ
Highlights
Stars
From Flatland to Space (SPAR). Accepted to NeurIPS 2025 Datasets & Benchmarks. A large-scale dataset & benchmark for 3D spatial perception and reasoning in VLMs.
🚀Official Repository of Intelligent Remote Sensing Agents: A Survey
Exclusively Dark (ExDARK) dataset which to the best of our knowledge, is the largest collection of low-light images taken in very low-light environments to twilight (i.e 10 different conditions) to…
[ICLR 2026] π^3: Permutation-Equivariant Visual Geometry Learning
[ICLR 2026] Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image
🚁 Can Vision-Language Models Think from the Sky? UAVReason for Aerial Reasoning and Generation
Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse
[ACL 2026] Code for "Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection"
[ACL 2026 Main] Official implementation of "Generating Attribution Reports for Manipulated Facial Images"
The official code of "Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search" [SIGIR 2026]
(ACM TOMM) This is the official code repository for "VM-UNet: Vision Mamba UNet for Medical Image Segmentation".
[PR 2026] Harnessing Weak Pair Uncertainty for Text-based Person Search
[IEEE Transactions on Image Processing'26] Pytorch implementation of FANet: Fovea Attention Network for Robust Aerial Geo-localization Across Diverse Weather Conditions
[ACL 2026] Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment
[ACL 2024 Findings] MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning https://arxiv.org/abs/2311.10537
The official implementation of Error Detection in Egocentric Procedural Task Videos
SUES-200: A Multi-height Multi-scene Cross-view Image Benchmark Across Drone and Satellite
[NeurIPS 2024] This repo contains evaluation code for the paper "Are We on the Right Way for Evaluating Large Vision-Language Models"
A Simple and Universal Swarm Intelligence Engine, Predicting Anything. 简洁通用的群体智能引擎,预测万物
TurtleBench: Evaluating Top Language Models via Real-World Yes/No Puzzles.
PyTorch building blocks for the OLMo ecosystem
Modeling, training, eval, and inference code for OLMo
Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
全维度的前沿大语言模型自动化评测套件。涵盖逻辑推理、智能体编程、网页特效代码生成以及百万Token级长文本解析(GPT-5.4 / Claude 4.7 / DeepSeek-V4 等)