Skip to content
View ZiyuGuo99's full-sized avatar
👋
👋

Block or report ZiyuGuo99

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

Python 89 2 Updated Aug 6, 2026

One Discrete Word for Visual Reasoning Overtakes Agentic and Latent Methods

Python 138 Updated Jun 9, 2026

The first Interleaved framework for textual reasoning within the visual generation process

Python 166 2 Updated Mar 16, 2026

Are Video Models Ready as Zero-shot Reasoners?

Python 87 4 Updated Nov 24, 2025

ULMEvalKit: One-Stop Eval ToolKit for Image Generation

Python 56 2 Updated Dec 17, 2025

A framework for unified personalized model, achieving mutual enhancement between personalized understanding and generation. Demonstrating the potential of cross-task information transfer in persona…

Python 131 2 Updated Jun 15, 2026

Official implementation of UnifiedReward & [NeurIPS 2025] UnifiedReward-Think & UnifiedReward-Flex

Python 799 41 Updated Jun 18, 2026

Official repository for the paper "TIIF-Bench: How Does Your T2I Model Follow Your Instructions?".

Python 125 4 Updated Jun 26, 2026

[NeurIPS 2025] MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning

Python 107 5 Updated Sep 19, 2025

Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning

Python 160 16 Updated Aug 1, 2025

EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights (CVPR 2025)

Python 14 1 Updated Apr 1, 2026

[NeurIPS 2025] T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Python 433 26 Updated Sep 18, 2025

MAGI-1: Autoregressive Video Generation at Scale

Python 3,760 238 Updated Jun 17, 2026

[CVPR 2025] The First Investigation of CoT Reasoning (RL, TTS, Reflection) in Image Generation

Python 863 28 Updated Mar 19, 2026

[ICLR 2025] The First Multimodal Seach Engine Pipeline and Benchmark for LMMs

Python 495 33 Updated Apr 5, 2026

The Most Faithful Implementation of Segment Anything (SAM) in 3D

Python 360 19 Updated Sep 11, 2024

[ECCV 2024] Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Python 183 17 Updated Apr 28, 2025

[CVPR 2024] OneLLM: One Framework to Align All Modalities with Language

Python 666 38 Updated Oct 22, 2024

[CVPR2024 Hightlight] No Time to Train: Empowering Non-Parametric Networks for Few-shot 3D Scene Segmentation

Python 118 6 Updated Apr 20, 2024

Align 3D Point Cloud with Multi-modalities for Large Language Models

Python 463 28 Updated Dec 9, 2023

Personalize Segment Anything Model (SAM) with 1 shot in 10 seconds

Python 1,671 112 Updated Jul 22, 2024
Python 388 76 Updated Feb 21, 2025

[ICCV 2023] Code for "Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior Refinement"

Jupyter Notebook 150 12 Updated Apr 21, 2024

(ICCV2023) Official implementation of 'ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding with GPT and Prototype Guidance'

C++ 60 5 Updated Apr 18, 2024

[ICLR 2024] Fine-tuning LLaMA to follow Instructions within 1 Hour and 1.2M Parameters

Python 5,912 378 Updated Mar 14, 2024

[CVPR 2023] Parameter is Not All You Need: Starting from Non-Parametric Networks for 3D Point Cloud Analysis

Python 542 50 Updated Apr 9, 2024

Official pytorch implementation of "DSPoint: Dual-scale Point Cloud Recognition with High-frequency Fusion"

Python 20 1 Updated Feb 4, 2025

Generic PyTorch dataset implementation to load and augment VIDEOS for deep learning training loops.

Python 471 45 Updated Jan 18, 2023

[CVPR 2023] Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners

Python 379 19 Updated Jun 1, 2023
Python 2 Updated Feb 2, 2024
Next