Highlights
- Pro
Stars
[ECCV 2026] AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation
Repository associated with paper titled "FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation", presented at ICML 2026.
A unified toolkit for efficient and effective coding agents (Karpathy principles, Caveman, Ponytail, RTK, CodeGraph, Context-Mode). Minimal setup under 30 seconds. Any OS.
[CVPR 2026 Highlight] official implementation of the paper: "FlowDC: Flow-Based Decoupling-Decay for Complex Image Editing"
Official inference repo for FLUX.1 models
A structured reading list on Vision-Language-Action (VLA) models β from diffusion/flow matching foundations through state-of-the-art robot foundation model architectures to data scaling, RL fine-tuβ¦
This is the homepage of a new book entitled "Mathematical Foundations of Reinforcement Learning."
NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.
A Survey on Reinforcement Learning of Vision-Language-Action Models for Robotic Manipulation
A general framework for inference-time scaling and steering of diffusion models with arbitrary rewards.
Theoretical introduction and examples for diffusion model algorithms
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex β developed and maintained with no human intervention.
Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
Resources needed to start deep learning research. ML/DL/CV/NLP/ML-SYS/RL/Graphs/Maths/Med image lecture videos from professors at esteemed universities.
Official Python inference and LoRA trainer package for the LTX-2 audioβvideo generative model.
Awesome Unified Multimodal Models
Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. π¦
[ACM MM 2025] SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
The official GitHub repo for the survey paper "A Survey on Diffusion Language Models".
AI agents running research on single-GPU nanochat training automatically
Official Code for "StreamCorrect: Bringing Offline Speech Recognition Performance to Streaming via Error Correction"
An OpenCode plugin that shows visual diffs in your editor after file edits.
Iterative Compositional Data Generation for Robot Control
[ICLR 2026] UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models
[ICLR2024] Official repo for paper "PnP Inversion: Boosting Diffusion-based Editing with 3 Lines of Code"
[ACM MM 2024] Training-free Cross-domain Image Composition via Adaptive Latent Manipulation and Energy-guided Optimization