Stars
[RSS 2025] CLIP-RT : Learning Language-Conditioned Robotic Policies from Natural Language Supervision
The official implementation of "Continuous SO(3) Equivariant Convolution for 3D Point Cloud Analysis" [ECCV 24]
This repo is meant to serve as a guide for Machine Learning/AI technical interviews.
Fine-Grained Causal Dynamics Learning with Quantization for Improving Robustness in Reinforcement Learning (ICML 2024)
[IROS 2024] PGA: Personalizing Grasping Agents with Single Human-Robot Interaction
🦾 PyTorch Implementation for the ICRA'24 Paper, "PROGrasp: Pragmatic Human-Robot Communication for Object Grasping"
Democratization of RT-2 "RT-2: New model translates vision and language into action"
[IROS 2023] GVCCI: Lifelong Learning of Visual Grounding for Language-Guided Robotic Manipulation
A curated list of reinforcement learning with human feedback resources (continually updated)
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
💬 Official PyTorch Implementation for CVPR'23 Paper, "The Dialog Must Go On: Improving Visual Dialog via Generative Self-Training"
SelecMix: Debiased Learning by Contradicting-pair Sampling (NeurIPS 2022)
Official implementation of History Aware Multimodal Transformer for Vision-and-Language Navigation (NeurIPS'21).
Code and Data of the CVPR 2022 paper: Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation
Official Implementation of IVLN-CE: Iterative Vision-and-Language Navigation in Continuous Environments
[CVPR 2022] Pseudo-Q: Generating Pseudo Language Queries for Visual Grounding
A curated list of Meta Learning papers, code, books, blogs, videos, datasets and other resources.
Audio Visual Scene-Aware Dialog (AVSD) Challenge at the 10th Dialog System Technology Challenge (DSTC)
Optimus: the first large-scale pre-trained VAE language model
PyTorch code for "Unifying Vision-and-Language Tasks via Text Generation" (ICML 2021)
Implementation for "Large-scale Pretraining for Visual Dialog" https://arxiv.org/abs/1912.02379
🌈 PyTorch Implementation for EMNLP'21 Findings "Reasoning Visual Dialog with Sparse Graph Learning and Knowledge Transfer"
Recent Advances in Vision and Language PreTrained Models (VL-PTMs)
Open-Source Toolkit for End-to-End Korean Automatic Speech Recognition leveraging PyTorch and Hydra.
We rank the 1st in DSTC8 Audio-Visual Scene-Aware Dialog competition. This is the source code for our IEEE/ACM TASLP (AAAI2020-DSTC8-AVSD) paper "Bridging Text and Video: A Universal Multimodal Tra…