Stars
State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
High-fidelity world models for general embodied intelligence, such as data engines and world simulators.
AI agents running research on single-GPU nanochat training automatically
[CVPR 2025] Official implementation of the paper "Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysis"
Robust 4D Action Segmentation in Corrupted Point Cloud Videos [ICME 2026]
Action Transition with Optimal Transport for Egocentric Procedural Error Detection [Accepted by ICME 2026]
Won the 25 ACM Multimedia Grand Challenge
[PG 2025] BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection via Real-Time Multi-View Box Fusion
A project page template for academic papers. Demo at https://eliahuhorwitz.github.io/Academic-project-page-template/
Official code release for ConceptGraphs
[CVPR 2025] Official codes for the paper 'Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models'
PyTorch implementation of Pointnet2/Pointnet++
This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!
g1: Using Llama-3.1 70b on Groq to create o1-like reasoning chains
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
800,000 step-level correctness labels on LLM solutions to MATH problems
📖 A curated list of resources dedicated to hallucination of multimodal large language models (MLLM).
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
[NeurIPS 2024] This repo contains evaluation code for the paper "Are We on the Right Way for Evaluating Large Vision-Language Models"
This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
[CVPR'24] HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models