Stars
[ICLR'26] Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
Official implementation of NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments (ICCV'25).
📖 This is a repository for organizing papers, codes, and other resources related to unified multimodal models.
[CVPR2025] Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters
[ICLR'26] SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models
Diffusion Models for Generative Outfit Recommendation
Personalized Image Generation with Large Multimodal Models
[KDD'25] Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective
[SIGIR'2024] "GraphGPT: Graph Instruction Tuning for Large Language Models"
A collection of resources on controllable generation with text-to-image diffusion models.