Stars
UAV World Model & VLA & VLM Learning - 无人机领域的世界模型、VLA、VLM综述学习项目
VLX-Flow: streaming VLM for real-time general vision intelligence
Awesome World Models for Aerial Navigation
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations https://video-prediction-policy.github.io
Official PyTorch Implementation of Unified Video Action Model (RSS 2025)
[RSS 2026] Causal video-action world model for generalist robot control
🚦 macOS menu bar traffic light for Claude Code — red (working), yellow (waiting for input), green (idle)
Building General-Purpose Robots Based on Embodied Foundation Model
A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
Markdown + 证件照 → 精美简历(PDF/HTML/PNG)| AI-powered resume generator from Markdown
WAM-Diff: A Masked Diffusion VLA Framework with MoE and Online Reinforcement Learning for Autonomous Driving
A curated, continuously updated reading list, paper blogs, and resources for World Action Models (WAMs) in embodied AI.
A Curated List of Vision-Language-Action (VLA) and World Action Models (WAM) Research and Beyond
[ECCV 2026] HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks
A list of research papers, models, datasets, and other resources on Vision-Language-Action/Navigation (VLA/VLN) models for UAVs.
ERNIE-Image is an open text-to-image generation model developed by the ERNIE-Image team at Baidu. It is built on a single-stream Diffusion Transformer (DiT), with only 8B DiT parameters, it reaches…
A rebuilt, fully functional version of Anthropic's Claude Code CLI
Corpus of Annotations for Misspelings