Lists (6)
Sort Name ascending (A-Z)
Stars
Efficient Triton Kernels for LLM Training
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
[🚀 ICLR 2026 Oral] NextStep-1: SOTA Autogressive Image Generation with Continuous Tokens. A research project developed by the StepFun’s Multimodal Intelligence team.
GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
An inference and training framework for multiple image input in Flux Kontext dev
Enjoy the magic of Diffusion models!
[TPAMI][ECCV2024 Oral] Clearer anytime frame interpolation & Manipulated interpolation of anything
A curated list of awesome research papers, projects, code, dataset, workshops etc. related to virtual try-on.
This repo contains two pixiv datasets: "colored-sketch pairs" and "popular pixiv images"
[IJCV2024] Exploiting Diffusion Prior for Real-World Image Super-Resolution
小红书(XiaoHongShu、RedNote)链接提取/作品采集工具:提取账号发布、收藏、点赞、专辑作品链接;提取搜索结果作品、用户链接;采集小红书作品信息;提取小红书作品下载地址;下载小红书作品文件
[NeurIPS 2025] Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
fork自BilibiliVideoDownload, 为了修复已知bug
Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.
Implementation of "YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception".
[ACL2026 Findings] GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
Train transformer language models with reinforcement learning.
Q-Insight Family: Q-Insight, VQ-Insight and RALI (NeurIPS 2025 Spotlight, AAAI 2026 Oral, and ICLR 2026 Oral)
[NeurIPS 2025] Image editing is worth a single LoRA! 0.1% training data for fantastic image editing! Surpasses GPT-4o in ID persistence~ MoE ckpt released! Only 4GB VRAM is enough to run!
A pipeline parallel training script for diffusion models.
Solve Visual Understanding with Reinforced VLMs
Code for Scaling Language-Free Visual Representation Learning (WebSSL)