Stars
Official repository for the paper "MICo-150K: A Comprehensive Dataset for Multi-Image Composition".
Qwen-Image text to image lora trainer
A curated collection of fun and creative examples generated with Nano Banana & Nano Banana Pro🍌, Gemini-2.5-flash-image based model. We also release Nano-consistent-150K openly to support the commu…
Repo for Qwen Image Finetune
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
[ICLR'26] Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
Enjoy the magic of Diffusion models!
A pipeline parallel training script for diffusion models.
Wan: Open and Advanced Large-Scale Video Generative Models
Official Implementation of Paper Transfer between Modalities with MetaQueries
A SOTA open-source image editing model, which aims to provide comparable performance against the closed-source models like GPT-4o and Gemini 2 Flash.
ACM MM 2023 paper: Semantic-based Selection Synthesis and Supervision for few-shot learning
[CVPR2025] Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters
[Official Implementation] Model Inversion Attacks through Target-specific Conditional Diffusion Models
[KDD'25] Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective
Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.
Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
This is an unofficial PyTorch implementation of StyleDrop: Text-to-Image Generation in Any Style.
Unoffical implement for [StyleDrop](https://arxiv.org/abs/2306.00983)
Reference implementation for DPO (Direct Preference Optimization)
《开源大模型食用指南》针对中国宝宝量身打造的基于Linux环境快速微调(全参数/Lora)、部署国内外开源大模型(LLM)/多模态大模型(MLLM)教程
LaVIT: Empower the Large Language Model to Understand and Generate Visual Content
[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.