Highlights
- Pro
Lists (9)
Sort Name ascending (A-Z)
Stars
F1: A Vision Language Action Model Bridging Understanding and Generation to Actions
[Tech Report] Context Scaling: Scaling Properties of Text Conditioning in Visual Generation
Eagle: Frontier Vision-Language Models with Data-Centric Strategies
Lightweight coding agent that runs in your terminal
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
We introduce BabyVision, a benchmark revealing the infancy of AI vision.
NEO Series: Native Vision-Language Models from First Principles
Companion code for the global workspace interpretability paper
[ICLR'25 Oral] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
Official implementation of Perceptive Behavior Foundation Model
PyTorch implementation of Parallel Rollout Approximation
[ICML 2025] CoreMatching: Co-adaptive Sparse Inference Framework for Comprehensive Acceleration of Vision Language Model
This is the official Python version of CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation.
[NeurIPS 2023]MathNAS: If Blocks Have a Role in Mathematical Architecture Design.
[ICLR 2025] Dobi-SVD : Differentiable SVD for LLM Compression and Some New Perspectives"
[NeurIPS Spotlight 2025] Angles Don’t Lie: Unlocking Training-Efficient RL Through the Model’s Own Signals.
JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System
📚 A curated collection of papers and open-source code repositories dedicated to the application of Vision-Language Models (VLMs) for streaming video.
Implementation of paper "Playful Agentic Robot Learning"
RoboBrain 2.5: Advanced version of RoboBrain. Depth in Sight, Time in Mind. 🎉🎉🎉
LAVIS - A One-stop Library for Language-Vision Intelligence
openvla / openvla
Forked from TRI-ML/prismatic-vlmsOpenVLA: An open-source vision-language-action model for robotic manipulation.
Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.
SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles