-
The Chinese University of Hong Kong
- Hong Kong
- https://ziyuguo99.github.io/
Stars
Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
One Discrete Word for Visual Reasoning Overtakes Agentic and Latent Methods
The first Interleaved framework for textual reasoning within the visual generation process
ULMEvalKit: One-Stop Eval ToolKit for Image Generation
A framework for unified personalized model, achieving mutual enhancement between personalized understanding and generation. Demonstrating the potential of cross-task information transfer in persona…
Official implementation of UnifiedReward & [NeurIPS 2025] UnifiedReward-Think & UnifiedReward-Flex
Official repository for the paper "TIIF-Bench: How Does Your T2I Model Follow Your Instructions?".
[NeurIPS 2025] MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights (CVPR 2025)
[NeurIPS 2025] T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
MAGI-1: Autoregressive Video Generation at Scale
[CVPR 2025] The First Investigation of CoT Reasoning (RL, TTS, Reflection) in Image Generation
[ICLR 2025] The First Multimodal Seach Engine Pipeline and Benchmark for LMMs
The Most Faithful Implementation of Segment Anything (SAM) in 3D
[ECCV 2024] Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
[CVPR 2024] OneLLM: One Framework to Align All Modalities with Language
[CVPR2024 Hightlight] No Time to Train: Empowering Non-Parametric Networks for Few-shot 3D Scene Segmentation
Align 3D Point Cloud with Multi-modalities for Large Language Models
Personalize Segment Anything Model (SAM) with 1 shot in 10 seconds
[ICCV 2023] Code for "Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior Refinement"
(ICCV2023) Official implementation of 'ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding with GPT and Prototype Guidance'
[ICLR 2024] Fine-tuning LLaMA to follow Instructions within 1 Hour and 1.2M Parameters
[CVPR 2023] Parameter is Not All You Need: Starting from Non-Parametric Networks for 3D Point Cloud Analysis
Official pytorch implementation of "DSPoint: Dual-scale Point Cloud Recognition with High-frequency Fusion"
Generic PyTorch dataset implementation to load and augment VIDEOS for deep learning training loops.
[CVPR 2023] Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners