Skip to content
View Wakals's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report Wakals

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

We introduce BabyVision, a benchmark revealing the infancy of AI vision.

Python 232 10 Updated Jan 13, 2026

NEO Series: Native Vision-Language Models from First Principles

Python 878 31 Updated Jul 1, 2026

Companion code for the global workspace interpretability paper

Python 1,578 230 Updated Jul 25, 2026

[ICLR'25 Oral] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Python 1,680 97 Updated Mar 16, 2025

Official implementation of Perceptive Behavior Foundation Model

Python 215 12 Updated Jul 8, 2026

PyTorch implementation of Parallel Rollout Approximation

Python 11 2 Updated Jun 29, 2026

[ICML 2025] CoreMatching: Co-adaptive Sparse Inference Framework for Comprehensive Acceleration of Vision Language Model

Python 16 2 Updated May 27, 2025

This is the official Python version of CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation.

Jupyter Notebook 18 2 Updated Oct 25, 2024

[NeurIPS 2023]MathNAS: If Blocks Have a Role in Mathematical Architecture Design.

Python 37 2 Updated Apr 10, 2024

[ICLR 2025] Dobi-SVD : Differentiable SVD for LLM Compression and Some New Perspectives"

Python 54 9 Updated Oct 19, 2025

[NeurIPS Spotlight 2025] Angles Don’t Lie: Unlocking Training-Efficient RL Through the Model’s Own Signals.

Python 83 10 Updated Sep 26, 2025

Code for the Molmo Vision-Language Model

Python 921 96 Updated Dec 12, 2024

JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System

Python 1,441 148 Updated Jul 22, 2026

Open-source unified multimodal model

Python 6,121 544 Updated May 4, 2026

📚 A curated collection of papers and open-source code repositories dedicated to the application of Vision-Language Models (VLMs) for streaming video.

189 5 Updated Jul 22, 2026

Implementation of paper "Playful Agentic Robot Learning"

Python 102 4 Updated Jun 20, 2026

RoboBrain 2.5: Advanced version of RoboBrain. Depth in Sight, Time in Mind. 🎉🎉🎉

Python 1,115 111 Updated Feb 28, 2026

LAVIS - A One-stop Library for Language-Vision Intelligence

Jupyter Notebook 11,258 1,109 Updated Jun 2, 2026

OpenVLA: An open-source vision-language-action model for robotic manipulation.

Python 6,702 805 Updated Mar 23, 2025

Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.

Python 2,103 117 Updated Jul 29, 2024

SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles

Python 4,398 378 Updated Jul 16, 2026

NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.

Jupyter Notebook 11,250 789 Updated Jul 24, 2026

[ICML 2026] RoboTwin 2.0 Offical Repo

Python 2,631 439 Updated Jul 25, 2026

Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals

Python 2,484 215 Updated Apr 19, 2026

Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷

Python 6,778 399 Updated Jul 24, 2026

A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.

196,533 20,232 Updated Apr 20, 2026

Fully Open Framework for Democratized Multimodal Training

Python 1,152 77 Updated Jul 26, 2026

Reinforcement Learning of Vision Language Models with Self Visual Perception Reward

Python 175 17 Updated Mar 14, 2026

A curated, continuously updated reading list, paper blogs, and resources for World Action Models (WAMs) in embodied AI.

HTML 1,172 30 Updated Jul 23, 2026
Next