-
Carnegie Mellon University
- Pittsburgh
- anuragxel.github.io
- @anuragxel
Stars
Research language for array processing in the Haskell/ML family
✨ An advanced 3D Gaussian Splatting renderer for THREE.js
Checkpoint and evaluation code for PhyCo [CVPR 2026]
GLUEMAP: Global Structure-from-Motion Meets Feedforward Reconstruction
PRIX: Learning to Plan from Raw Pixels for End-to-End Autonomous Driving
[ECCV 2026] FrameCrafter: Novel View Synthesis as Video Completion
[CVPR 2025] StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
Implementation of Open-World Visual Odometry with Temporal Dynamics Awareness (CVPR'26)
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
DSPy: The framework for programming—not prompting—language models
123D: A Unified Library for Multi-Modal Autonomous Driving Data
ClickHouse® is a real-time analytics database management system
COLMAP - Structure-from-Motion and Multi-View Stereo
A repository for the multimodal WorkZone3D dataset
Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation
Fail2Drive: Benchmarking Closed-Loop Driving Generalization
Allen Institute for AI: WildDet3D: Scaling Promptable 3D Detection in the Wild
A JAX-based simulator for autonomous driving research.
[ECCV 2026 Oral] LoMa: Local Feature Matching Revisited
A optimized PyTorch framework for behavior cloning with flow related generative models.
[CVPR 2025] Multiple Object Tracking as ID Prediction
PyTorch building blocks for the OLMo ecosystem
Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
The repository provides code for running inference with the Meta Segment Anything Audio Model (SAM-Audio), links for downloading the trained model checkpoints, and example notebooks that show how t…
Cosmos-Transfer2.5, built on top of Cosmos-Predict2.5, produces high-quality world simulations conditioned on multiple spatial control inputs.