Hi, I'm Mingqiang Tang
I’m an undergraduate student majoring in Computer Science and Technology at the Southern University of Science and Technology (SUSTech), Shenzhen, China. My current research interests are multimodal learning and multimodal large language models.
Education
Southern University of Science and Technology (SUSTech) · 2023 – Expected 2027
Bachelor of Engineering in Computer Science and Technology · Shenzhen, China
- GPA: 3.86 / 4.00, Rank: 16 / 152
| Course | Grade |
|---|---|
| Computer Vision | A+ |
| Design and Analysis of Algorithms | A+ |
| Advanced Programming | A+ |
| Data Structures | A |
| Discrete Mathematics | A |
| Digital Logic | A |
| Introduction to Mathematical Logic | A |
| Computer Networks | A |
| Natural Language Processing | A |
| Software Engineering | A |
Research
Multimodal LLM-based Multi-turn Composed Image Retrieval · Feb. 2026 – Jun. 2026 · arXiv:2607.20291
SUSTech; advised by Prof. Xuemeng Song · Shenzhen, China
- Surveyed and reproduced related multi-turn composed image retrieval methods, analyzed the limitations of existing work, and extended the task to scenarios supporting multi-intent switching.
- Built a multi-turn composed image retrieval dataset from public single-turn fashion retrieval datasets via cross-dataset visual bridging, supporting multimodal query forms (text, image, sketch, and image-text combinations) and covering diverse fashion retrieval tasks.
- Designed a model that maps multi-turn context into a unified retrieval space via background-removed image alignment, fashion image-text alignment, and LoRA fine-tuning of a multimodal LLM, enabling target image retrieval.
- Completed the paper as first author; it is currently under review at a CCF-A conference.
LLM-based Automated Unit Test Generation for Rust Programs · Jun. 2025 – Jan. 2026
Research Intern at SQLab, SUSTech; advised by Prof. YePang Liu; collaborated with Ant Group · Shenzhen, China
- Built an evaluation dataset for automated Rust unit test generation, covering typical program bugs and test scenarios from open-source projects (GitHub).
- Experimentally evaluated Rust test generation tools including RPG, RustyUnit, and RUG, comparing their code coverage, bug-triggering capability, and test effectiveness across multiple project types (GitHub).
- Analyzed the causes of compilation failures and insufficient coverage when LLMs generate Rust unit test code, summarizing key issues in type inference, dependency resolution, API call constraints, and code executability.
Projects
Structured Image Generation with Spatial Reasoning Models
Uncertainty-Aware Structured Image Generation · GitHub
- Investigated the discrete constraint reasoning capabilities of denoising generative models in continuous image spaces using SRM, and implemented uncertainty-driven adaptive sequential generation.
- Designed three inference strategies—dynamic patch grouping, uncertainty-guided local re-noising, and learned online backtracking—to improve generation efficiency and global structural consistency, and analyzed how parallel generation, local correction, and constraint feedback affect complex compositional reasoning.
Sketch-and-Text Composed Image Retrieval
TASK-former-based Composed Image Retrieval · GitHub
- Trained and evaluated a sketch-and-text composed image retrieval model based on TASK-former, supporting text-only, sketch-only, and sketch-text combined queries.
- Conducted optimization experiments on dynamic feature fusion, stroke-level data augmentation, and dual-view feature consistency to bridge the domain gap between synthetic and real hand-drawn sketches and improve generalization to real sketches.
CLIP and LLM2CLIP Semantic Perturbation Sensitivity Analysis · May 2026
Semantic Sensitivity Analysis of Vision-Language Models · GitHub
- Constructed three types of controlled semantic perturbations on the MSCOCO dataset—object replacement, color/spatial-relationship replacement, and semantic interference—and quantified the semantic sensitivity of CLIP and LLM2CLIP via the change in image-text similarity before and after perturbation.
- Implemented a complete pipeline for batch inference evaluation, result statistics, and visualization, comparing the fine-grained semantic understanding of the models across tens of thousands of image-text pairs.
Multi-turn Fashion Image Retrieval · Mar. 2026
FastAPI-based Multi-turn Composed Image Retrieval System · GitHub · Demo
- Built a multi-turn fashion image retrieval system with FastAPI and the CFIR model, enabling users to progressively refine their retrieval intent through multi-turn image and text feedback.
- Implemented model inference, vector retrieval, index caching, and Top-K retrieval result display and selection, completing a full multi-turn interactive image retrieval workflow.
Assignment Management System · Feb. 2025 – Jun. 2025
Full-stack Teaching and Assignment Management Platform · GitHub
- Collaboratively developed an intelligent assignment management platform for university teaching, with a decoupled frontend-backend architecture based on Spring Boot and TypeScript.
- Integrated OCR document recognition, Judge0 code evaluation, and AI-assisted grading to support automated grading of document and code assignments; implemented core workflows including role-based permission control, class management, and assignment release and submission.
Awards
- SUSTech Merit Student Scholarship – Second Class · SUSTech · 2024
- Outstanding Student · SUSTech · 2024, 2025
- SUSTech Merit Student Scholarship – Third Class · SUSTech · 2025
Technical Skills
- Programming: Java, Rust, C/C++, Python
- Development: Linux, Spring Boot, PostgreSQL
- Language: CET-6: 492; CET Spoken English Test: Good
Interests
- Sports: Swimming, Running, Frisbee
- Volunteering: 73.5 hours of volunteer service accumulated
