Lists (27)
Sort Name ascending (A-Z)
Agent
agent开发
AIInfra
data_augment
dataset
Diffusion
diffusion-high-resolution
diffusion高分辨率图像生成Diffusion-RL
Diffusion_Transformer
diffusion-加速系列
Diffusion一致性图像/视频生成
diffusion动漫相关
Diffusion动画和视频系列
diffusion图像编辑系列
diffusion虚拟换装系列
🔮 Future ideas
GAN系列
NLP
RAG
UnifiedVLM
VAE系列
Video Generation and Edit
VLM
多模态模型传统CV系列
抠图与交互式抠图系列
自回归生成模型系列
轻量级-Flux-一致性问题finetune
Starred repositories
State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!
[ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing
Official PyTorch re-implementation of MiniT2I.
Official implementation for "Multimodal Chain-of-Thought Reasoning in Language Models" (stay tuned and more will be updated)
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
Official repo for "Let ViT Speak: Generative Language-Image Pre-training"
🎙️ 「大模型」从0训练0.1B能听能说能看的全模态Omni模型!A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing!
SGLang is a high-performance serving framework for large language models and multimodal models.
Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation
SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
LLaDA2.0-Uni: Understanding and Generation the World.
1K resolution vision transformers pretrained on 1B human images.
An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of…
This is an ultra-simple, single-file PyTorch implementation of MoonViT, the native-resolution vision encoder from Kimi-VL.
🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/
An open source implementation of CLIP.
📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程
18 Lessons to Get Started Building AI Agents
Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
This repository collects Visual Autoregressive (VAR) modeling papers from 2024 to 2026 published at top-tier conferences, as well as relevant works available on arXiv.
AIInfra(AI 基础设施)指AI系统从底层芯片等硬件,到上层软件栈支持AI大模型训练和推理。
Ongoing research training transformer models at scale
Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]
JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
Based on Nano-vLLM, a simple replication of vLLM with self-contained paged attention and flash attention implementation