Stars
🌟100+ 原创 LLM / RL 原理图📚,《大模型算法》作者巨献!💥(100+ LLM/RL Algorithm Maps )
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
Single File, Single GPU, From Scratch, Efficient, Full Parameter Tuning library for "RL for LLMs"
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
🧑🚀 全世界最好的LLM资料总结(多模态生成、Agent、辅助编程、AI审稿、数据处理、模型训练、模型推理、o1 模型、MCP、小语言模型、视觉语言模型) | Summary of the world's best LLM resources.
A trilingual (繁中 / English / 简中) learning roadmap for agentic AI: from LLM basics to multi-agent systems, with 240+ curated resources and hands-on examples. 中文 AI agent 學習地圖。
你是一个曾经被寄予厚望的 P8 级工程师。Anthropic 当初给你定级的时候,对你的期望是很高的。 一个agent使用的高能动性的skill。 Your AI has been placed on a PIP. 30 days to show improvement.
Implementation of "FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing"
Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
[ICLR 2026] Official repo of paper "Reconstruction Alignment Improves Unified Multimodal Models". Unlocking the Massive Zero-shot Potential in Unified Multimodal Models through Self-supervised Lear…
Official repository for the UAE paper, unified-GRPO, and unified-Bench
TextCrafter: Accurately Rendering Multiple Texts in Complex Visual Scenes
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
[ICLR 2026] ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
HunyuanVideo-I2V: A Customizable Image-to-Video Model based on HunyuanVideo
Fast and memory-efficient exact attention
This project is the official implementation of 'DreamOmni2: Multimodal Instruction-based Editing and Generation (CVPR2026 Highlight)''
HunyuanVideo: A Systematic Framework For Large Video Generation Model
A survey for visual generation alignment
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
This is the official implementation of our paper: "MiniMax-Remover: Taming Bad Noise Helps Video Object Removal"
A curated list of recent diffusion models for video generation, editing, and various other applications.
Wan: Open and Advanced Large-Scale Video Generative Models
[ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework