Stars
Benchmark to measure what the real knowledge cutoff of a model is
[ICML'26] Code and website for Self-Flow: Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
[ICLR2026] The official code of "Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance"
A curated list of papers and selected technical blogs on Loop Models.
A theoretical reconstruction of the Claude Mythos architecture, built from first principles using the available research literature.
The best OSS video generation models, created by Genmo
[NeurIPS 2025 Oral] Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
🟣 LLMs interview questions and answers to help you prepare for your next machine learning and data science interview in 2026.
🔥 LeetCode for PyTorch — practice implementing softmax, attention, GPT-2 and more from scratch with instant auto-grading. Jupyter-based, self-hosted or try online.
Autonomous AI movie studio — turn a text prompt into a fully produced video. 100% local, no cloud, no API keys.
Official release of the benchmark in paper "VSP: Diagnosing the Dual Challenges of Perception and Reasoning in Spatial Planning Tasks for MLLMs"
Minimal and highly hackable implementation of Looped Transformers with GPT
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data (NeurIPS 2023 Spotlight) / / / / When Does Perceptual Alignment Benefit Vision Representations? (NeurIPS 2024)
[ICLR 2025] MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
Curated list of methods that focuses on improving the efficiency of diffusion models
A PyTorch implementation of the paper "All are Worth Words: A ViT Backbone for Diffusion Models".
https://www.shoufachen.com/Awesome-Diffusion-Transformers/
A PyTorch implementation of the paper "Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis"
Minimal reproduction of DeepSeek R1-Zero
Denoising Diffusion Step-aware Models (ICLR2024)
[ICML 2024] CLLMs: Consistency Large Language Models
[CVPR2025 Highlight] PAR: Parallelized Autoregressive Visual Generation. https://yuqingwang1029.github.io/PAR-project
Multi-Agent System Powered by LLMs for End-to-end Multimodal ML Automation
[CVPR 2024] DeepCache: Accelerating Diffusion Models for Free
Official Implementation for our NeurIPS 2024 paper, "Don't Look Twice: Run-Length Tokenization for Faster Video Transformers".
Official codebase used to develop Vision Transformer, SigLIP, MLP-Mixer, LiT and more.
High-Resolution Image Synthesis with Latent Diffusion Models