Stars
[ECCV 2026] The official implementation of paper "PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion"
[ECCV 2026] Official implementation of Hi-DiT: Hybrid Latent-Pixel Diffusion Transformer for Image Generation
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
Ideogram 4: Open image model at the forefront of design
Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation
SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
[NeurIPS 2025] Official implementation for our paper "Scaling Diffusion Transformers Efficiently via μP".
Train the smallest LM you can that fits in 16MB. Best model wins!
Fast, accurate & comprehensive text measurement & layout
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
AI agents running research on single-GPU nanochat training automatically
[ICML'26] Code and website for Self-Flow: Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
NextFlow🚀: Unified Sequential Modeling Activates Multimodal Understanding and Generation
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
A research-friendly PyTorch Lightning toolkit for training, fine-tuning, and evaluating AutoencoderKL for Stable Diffusion and FLUX.
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
[ICLR 2026] Official repo of paper "Reconstruction Alignment Improves Unified Multimodal Models". Unlocking the Massive Zero-shot Potential in Unified Multimodal Models through Self-supervised Lear…
Ming - facilitating advanced multimodal understanding and generation capabilities built upon the Ling LLM.
[NeurIPS 2025] Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
The official repository for LaCTok:Latent Consistency Tokenizer for High-resolution Image Reconstruction and Generation by 256 Tokens
Enjoy the magic of Diffusion models!
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo