Stars
[RSS 2025] CLIP-RT : Learning Language-Conditioned Robotic Policies from Natural Language Supervision
The official implementation of "Continuous SO(3) Equivariant Convolution for 3D Point Cloud Analysis" [ECCV 24]
This repo is meant to serve as a guide for Machine Learning/AI technical interviews.
openvla / openvla
Forked from TRI-ML/prismatic-vlmsOpenVLA: An open-source vision-language-action model for robotic manipulation.
Fine-Grained Causal Dynamics Learning with Quantization for Improving Robustness in Reinforcement Learning (ICML 2024)
[IROS 2024] PGA: Personalizing Grasping Agents with Single Human-Robot Interaction
🦾 PyTorch Implementation for the ICRA'24 Paper, "PROGrasp: Pragmatic Human-Robot Communication for Object Grasping"
Democratization of RT-2 "RT-2: New model translates vision and language into action"
[IROS 2023] GVCCI: Lifelong Learning of Visual Grounding for Language-Guided Robotic Manipulation
A curated list of reinforcement learning with human feedback resources (continually updated)
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
💬 Official PyTorch Implementation for CVPR'23 Paper, "The Dialog Must Go On: Improving Visual Dialog via Generative Self-Training"
SelecMix: Debiased Learning by Contradicting-pair Sampling (NeurIPS 2022)
Official implementation of History Aware Multimodal Transformer for Vision-and-Language Navigation (NeurIPS'21).
Code and Data of the CVPR 2022 paper: Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation
Official Implementation of IVLN-CE: Iterative Vision-and-Language Navigation in Continuous Environments
[CVPR 2022] Pseudo-Q: Generating Pseudo Language Queries for Visual Grounding
A curated list of Meta Learning papers, code, books, blogs, videos, datasets and other resources.
Code for SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations
Audio Visual Scene-Aware Dialog (AVSD) Challenge at the 10th Dialog System Technology Challenge (DSTC)
Optimus: the first large-scale pre-trained VAE language model
PyTorch code for "Unifying Vision-and-Language Tasks via Text Generation" (ICML 2021)
Implementation for "Large-scale Pretraining for Visual Dialog" https://arxiv.org/abs/1912.02379
🌈 PyTorch Implementation for EMNLP'21 Findings "Reasoning Visual Dialog with Sparse Graph Learning and Knowledge Transfer"
Recent Advances in Vision and Language PreTrained Models (VL-PTMs)
Conceptual 12M is a dataset containing (image-URL, caption) pairs collected for vision-and-language pre-training.