Highlights
- Pro
Stars
A generalist video MLLM built for fine-grained motion, long-form reasoning, temporal grounding, and online proactive response.
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
Evaluation code for "Benchmarking Visual State Tracking in Multimodal Video Understanding"
Official implementation of EgoHOD at ICLR 2025; 14 EgoVis Challenge Winners in CVPR 2024
Official implementation of "HowToCaption: Prompting LLMs to Transform Video Annotations at Scale." ECCV 2024
[ICCV2023] EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object Understanding
[Nips 2025] EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
TIPSv2 (CVPR'26) and TIPS (ICLR'25)
[CVPR 2026 Highlight] A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
[ICML 2026] Official implementation of "Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression"
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
Frontier Multimodal Foundation Models for Image and Video Understanding
Accepted By The 39th Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track
[NeurIPS 2025 Oral]Infinity⭐️: Unified Spacetime AutoRegressive Modeling for Visual Generation
Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance (CVPR 2026 Highlight)
[RSS 2023] Diffusion Policy Visuomotor Policy Learning via Action Diffusion
MotionStream: Real-Time Video Generation with Interactive Motion Controls
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
Official PyTorch implementation of "Video Summarization with Large Language Models" (CVPR 2025).
Native Multimodal Models are World Learners
[ICML2025] The code and data of Paper: Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
Official Repo for Self-Forcing++ High Quality Long Video Generation
Video Generation, Physical Commonsense, Semantic Adherence, VideoCon-Physics