-
horizon robotics
- nanjing, china
Stars
Make any agent harness multimodal-native.
An LLM post-training framework with vLLM for RL Scaling
A Curated List of Vision-Language-Action (VLA) and World Action Models (WAM) Research and Beyond
Official repository for "CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation"
Training Large Language Model to Reason in a Continuous Latent Space
(ICML2026) Official implementation of VLANeXt.
[ECCV 2026] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
A curated list of awesome LLM/VLM/VLA/World Model for Autonomous Driving(LLM4AD) resources (continually updated)
AutoGaze automatically removes redundant patches in a video, reducing #tokens in ViT/MLLM by 4x-100x.
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
1K resolution vision transformers pretrained on 1B human images.
LLaDA2.0-Uni: Understanding and Generation the World.
Self-evolving vision language models from zero data
This repository contains the code to train and evaluate TRIBE v2, a multimodal model for brain response prediction
Qwen3.6 is the large language model series developed by Qwen team, Alibaba Group.
[CVPR2026] VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
Awesome Unified Multimodal Models
"RAG-Anything: All-in-One RAG Framework"
Object detection on multiple datasets with an automatically learned unified label space.
[CVPR2026] Detect Anything via Next Point Prediction
LLM2CLIP significantly improves already state-of-the-art CLIP models.
Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"
[NeurIPS 2025] Official code for JAFAR: Jack up Any Feature at Any Resolution