Stars
An easy-to-use, fast toolkit to scale up RL post-training on a single node.
Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022
DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis
TAT-QA (Tabular And Textual dataset for Question Answering) contains 16,552 questions associated with 2,757 hybrid contexts from real-world financial reports.
Skills for Real Engineers. Straight from my .agents directory.
The code and resource of "Towards Comprehensive Detection of Chinese Harmful Memes" (NeurIPS2024 D&B).
[AAAI'25 (Oral)] Jailbreaking Large Vision-language Models via Typographic Visual Prompts
[ACL 2025] Can We Trust AI Doctors? A Survey of Medical Hallucination in Large Language and Large Vision-Language Models
up-to-date curated list of state-of-the-art Large vision language models hallucinations research work, papers & resources
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
An agentic skills framework & software development methodology that works.
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
A toy PyTorch implementation of FLUX diffusion transformers
斯坦福CS146S 现代软件开发者(vibe coding)课程中文版。 本中文课程由RapidAI 赞助。英文版: https://themodernsoftware.dev/
[ICLR'26] SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models
Algorithm powering the For You feed on X
Official implementation of NeurIPS'24 paper "Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models". This work adversarially unlearns the text encoder to enh…
Image inpainting tool powered by SOTA AI Model. Remove any unwanted object, defect, people from your pictures or erase and replace(powered by stable diffusion) any thing on your pictures.
Official Repo for Paper "OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision" [ICLR2025]
The image prompt adapter is designed to enable a pretrained text-to-image diffusion model to generate images with image prompt.
The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
[ICCV 2023] Consistent Image Synthesis and Editing
Officail Implementation for "ReNoise: Real Image Inversion Through Iterative Noising"