Skip to content
View YiyanXu's full-sized avatar

Block or report YiyanXu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Official repository for the paper "MICo-150K: A Comprehensive Dataset for Multi-Image Composition".

Python 104 2 Updated Apr 21, 2026

Qwen-Image text to image lora trainer

Python 766 70 Updated Dec 16, 2025

A curated collection of fun and creative examples generated with Nano Banana & Nano Banana Pro🍌, Gemini-2.5-flash-image based model. We also release Nano-consistent-150K openly to support the commu…

23,541 2,390 Updated Dec 12, 2025

Repo for Qwen Image Finetune

Jupyter Notebook 253 27 Updated Aug 11, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 23,049 4,431 Updated Aug 20, 2026

[ICLR'26] Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?

Python 55 1 Updated Mar 9, 2026

Open-source SOTA multi-image editing model

Python 871 43 Updated Jul 13, 2026

Enjoy the magic of Diffusion models!

Python 12,973 1,273 Updated Aug 20, 2026

A pipeline parallel training script for diffusion models.

Python 2,012 285 Updated Aug 20, 2026

Wan: Open and Advanced Large-Scale Video Generative Models

Python 17,222 2,179 Updated Mar 17, 2026

Official Implementation of Paper Transfer between Modalities with MetaQueries

Python 326 14 Updated Oct 12, 2025
JavaScript 4,332 1,939 Updated Jun 21, 2024

A SOTA open-source image editing model, which aims to provide comparable performance against the closed-source models like GPT-4o and Gemini 2 Flash.

Python 2,253 101 Updated Apr 29, 2026
Python 13 Updated Feb 25, 2025

ACM MM 2023 paper: Semantic-based Selection Synthesis and Supervision for few-shot learning

Python 4 Updated Nov 28, 2023

[CVPR2025] Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters

Jupyter Notebook 44 Updated Mar 11, 2025

[Official Implementation] Model Inversion Attacks through Target-specific Conditional Diffusion Models

Python 11 4 Updated Sep 7, 2025

[KDD'25] Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective

Python 113 18 Updated Jun 30, 2026

Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.

Python 2,103 118 Updated Jul 29, 2024

Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI

Python 1,390 75 Updated Aug 20, 2026

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Python 326 14 Updated Apr 17, 2025

This is an unofficial PyTorch implementation of StyleDrop: Text-to-Image Generation in Any Style.

Python 226 15 Updated Jul 11, 2023

Unoffical implement for [StyleDrop](https://arxiv.org/abs/2306.00983)

Python 587 30 Updated Aug 23, 2023

Reference implementation for DPO (Direct Preference Optimization)

Python 2,906 236 Updated Aug 11, 2024

《开源大模型食用指南》针对中国宝宝量身打造的基于Linux环境快速微调(全参数/Lora)、部署国内外开源大模型(LLM)/多模态大模型(MLLM)教程

Jupyter Notebook 31,787 3,089 Updated Jul 30, 2026
Python 4 Updated Jan 15, 2026

LaVIT: Empower the Large Language Model to Understand and Generate Visual Content

Jupyter Notebook 604 33 Updated Oct 6, 2024

[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)

Jupyter Notebook 1,842 100 Updated Feb 1, 2025

[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.

Python 24,992 2,778 Updated Aug 12, 2024
Next