Stars
Repository associated with paper titled "CLP: Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think"
Repository associated with paper titled "RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis"
Training non-differentiable networks via optimal transport.
🔥 A curated roadmap to the Efficient VLA landscape. We’re keeping this list live—contribute your latest work!
Repository associated with paper titled "TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching", presented at ICML 2026.
Repository associated with paper titled "FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation", presented at ICML 2026.
A Curated List of Vision-Language-Action (VLA) and World Action Models (WAM) Research and Beyond
Training-free structured sparsification for SAM.
[ICLR_2026] A Scalable Fragment-Aware Conformer Ensemble Transformer
OBEYED-VLA: Clutter-Resistant Vision-Language-Action Models through Object-Centric and Geometry Grounding
Official Implementation of "Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning." ICLR 2026.
[TMLR] DuFal: Dual-Frequency-Aware Learning for High-Fidelity Extremely Sparse-view CBCT Reconstruction
[ICLR 2024] MARCEL: Machine Learning over Molecular Conformer Ensembles
[NeurIPS 2025] How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?
S-Chain: Structured Visual Chain-of-Thought For Medicine
[ICML2025, NeurIPS2025 Spotlight] Sparse VideoGen 1 & 2: Accelerating Video Diffusion Transformers with Sparse Attention
[NeurIPS 2025] ExGra-Med: Medical Multi-Modal LLM with Extended Context Alignment
lecture slides for Deepmind x UCL 2021 reinforcement learning course available in YouTbue
Official implementation of Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation. Accepted in NeurIPS 2025.
Hierarchical Refinement: Optimal Transport to Infinity and Beyond
JAX - A curated list of resources https://github.com/google/jax
Official implementation of "OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning"