Highlights
- Pro
Lists (9)
Sort Name ascending (A-Z)
Stars
We introduce BabyVision, a benchmark revealing the infancy of AI vision.
NEO Series: Native Vision-Language Models from First Principles
Companion code for the global workspace interpretability paper
[ICLR'25 Oral] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
Official implementation of Perceptive Behavior Foundation Model
PyTorch implementation of Parallel Rollout Approximation
[ICML 2025] CoreMatching: Co-adaptive Sparse Inference Framework for Comprehensive Acceleration of Vision Language Model
This is the official Python version of CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation.
[NeurIPS 2023]MathNAS: If Blocks Have a Role in Mathematical Architecture Design.
[ICLR 2025] Dobi-SVD : Differentiable SVD for LLM Compression and Some New Perspectives"
[NeurIPS Spotlight 2025] Angles Don’t Lie: Unlocking Training-Efficient RL Through the Model’s Own Signals.
JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System
📚 A curated collection of papers and open-source code repositories dedicated to the application of Vision-Language Models (VLMs) for streaming video.
Implementation of paper "Playful Agentic Robot Learning"
RoboBrain 2.5: Advanced version of RoboBrain. Depth in Sight, Time in Mind. 🎉🎉🎉
LAVIS - A One-stop Library for Language-Vision Intelligence
openvla / openvla
Forked from TRI-ML/prismatic-vlmsOpenVLA: An open-source vision-language-action model for robotic manipulation.
Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.
SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
[ICML 2026] RoboTwin 2.0 Offical Repo
Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
Fully Open Framework for Democratized Multimodal Training
Reinforcement Learning of Vision Language Models with Self Visual Perception Reward
A curated, continuously updated reading list, paper blogs, and resources for World Action Models (WAMs) in embodied AI.