Skip to content
View laserwave's full-sized avatar
  • horizon robotics
  • nanjing, china

Block or report laserwave

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Make any agent harness multimodal-native.

Python 2,513 137 Updated Aug 13, 2026

An LLM post-training framework with vLLM for RL Scaling

Python 419 77 Updated Aug 3, 2026

A Curated List of Vision-Language-Action (VLA) and World Action Models (WAM) Research and Beyond

964 36 Updated Aug 8, 2026

Official repository for "CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation"

Python 104 20 Updated Dec 15, 2025

Training Large Language Model to Reason in a Continuous Latent Space

Python 1,683 187 Updated Jul 2, 2026

(ICML2026) Official implementation of VLANeXt.

Python 222 10 Updated Aug 10, 2026

[ECCV 2026] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model

Python 537 47 Updated May 2, 2026

NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.

Jupyter Notebook 11,511 826 Updated Aug 14, 2026

SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer

Python 8,766 700 Updated Aug 14, 2026

A curated list of awesome LLM/VLM/VLA/World Model for Autonomous Driving(LLM4AD) resources (continually updated)

1,889 111 Updated Jun 22, 2026
Python 466 51 Updated May 28, 2026

AutoGaze automatically removes redundant patches in a video, reducing #tokens in ViT/MLLM by 4x-100x.

Python 301 24 Updated May 5, 2026

StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

Python 3,459 448 Updated Aug 9, 2026

SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles

Python 4,790 409 Updated Aug 13, 2026

1K resolution vision transformers pretrained on 1B human images.

Python 904 61 Updated May 24, 2026

LLaDA2.0-Uni: Understanding and Generation the World.

Python 724 46 Updated May 29, 2026

Self-evolving vision language models from zero data

Python 81 2 Updated Mar 14, 2026

This repository contains the code to train and evaluate TRIBE v2, a multimodal model for brain response prediction

Jupyter Notebook 3,132 689 Updated Jun 23, 2026

Qwen3.6 is the large language model series developed by Qwen team, Alibaba Group.

3,791 266 Updated Jun 3, 2026

[CVPR2026] VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice

Python 88 5 Updated Feb 27, 2026

DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference

Python 25 1 Updated May 21, 2026

Awesome Unified Multimodal Models

1,310 46 Updated Mar 24, 2026

Official implementation of BLIP3o-Series

Python 1,665 79 Updated Nov 29, 2025

"RAG-Anything: All-in-One RAG Framework"

Python 22,913 2,658 Updated Aug 13, 2026

Grounded Language-Image Pre-training

Python 2,606 217 Updated Jan 24, 2024

Object detection on multiple datasets with an automatically learned unified label space.

Python 517 56 Updated Mar 8, 2024

[CVPR2026] Detect Anything via Next Point Prediction

Jupyter Notebook 1,548 111 Updated Feb 22, 2026

LLM2CLIP significantly improves already state-of-the-art CLIP models.

Python 683 32 Updated Feb 1, 2026

Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"

Python 1,990 87 Updated Feb 25, 2026

[NeurIPS 2025] Official code for JAFAR: Jack up Any Feature at Any Resolution

Jupyter Notebook 239 15 Updated Nov 24, 2025
Next