Skip to content
View gf0507033's full-sized avatar

Block or report gf0507033

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

👀「大模型」2小时从0训练65M参数的视觉多模态VLM!Train a 65M-parameter VLM from scratch in just 2h!

Python 8,384 920 Updated Jun 28, 2026

UFM: A Unified Dense Image Correspondence Estimator for both Optical Flow & Wide Baseline Matching Tasks. Matches any pair of images. (NeurIPS 2025)

Python 340 22 Updated Apr 4, 2026

One framework to evaluate any VLA model on any robot simulation benchmark.

Python 489 37 Updated Jul 29, 2026

VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo

Python 2,112 239 Updated Jul 29, 2026

ACL 2026 Main

Python 13 Updated Jul 15, 2026

[ICML 2026] ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning

Python 81 1 Updated Jul 8, 2026

1K resolution vision transformers pretrained on 1B human images.

Python 884 61 Updated May 24, 2026

INF Tech's open-source MLLMs for SOTA visual-language understanding and advanced document intelligence.

Python 239 25 Updated Jul 22, 2026

Allen Institute for AI: WildDet3D: Scaling Promptable 3D Detection in the Wild

Python 600 44 Updated Jun 1, 2026

SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments

Python 80 2 Updated Apr 16, 2026

This is a repository for listing papers on scene graph generation and application.

705 49 Updated Jul 10, 2026

This repository collects and organises state‑of‑the‑art papers on spatial reasoning for Multimodal Vision–Language Models (MVLMs).

319 17 Updated Feb 17, 2026

MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.

Python 5,266 687 Updated Jul 28, 2026

Обход блокировок рунета 2026

256 10 Updated Jun 6, 2026

Build scalable data pipelines on YTsaurus with automatic stage management, local development simulation, and more.

Python 28 Updated Jul 23, 2026

NVIDIA Alpamayo 1.5 Nano is an open 10B reasoning VLA model for autonomous vehicles with reinforcement-learning enhanced reasoning, navigation guidance, and visual question answering.

Python 345 66 Updated Jul 22, 2026

[CVPR26] LEAD: Minimizing Learner–Expert Asymmetry in End-to-End Driving

Python 208 24 Updated Jul 22, 2026

NVIDIA Alpamayo 1 Nano is an open 10B reasoning VLA model for autonomous vehicles that pairs driving trajectories with Chain-of-Causation reasoning.

Python 1,949 325 Updated Jul 22, 2026

AlpaSim is an open-source autonomous vehicle simulation platform designed for development and testing of end-to-end AV policies

Python 1,144 143 Updated Jul 21, 2026

Official implementation of Kimodo, a kinematic motion diffusion model for high-quality human(oid) motion generation.

Python 3,042 339 Updated Jul 13, 2026

Seoul World Model: Grounding World Simulation Models in a Real-World Metropolis

621 21 Updated Mar 17, 2026

The ultimate collection of high-fidelity Seedance 2.0 prompts and Seedance AI resources. Discover Seedance 2.0 how to use for cinematic film, anime, UGC, social media, meme and advertising. Include…

Shell 2,224 248 Updated Jul 29, 2026

[ICML 2026 Oral] Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence

Python 372 8 Updated Jul 26, 2026

[CVPR 2026] Offical implementation of the paper "HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images".

Python 106 5 Updated Jun 7, 2026

Machine Learning Systems

Python 27,629 3,380 Updated Jul 28, 2026

[ECCV 2026] Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels

Python 259 23 Updated Jul 3, 2026

[ICML'26] Official repository of Utonia: Toward One Encoder for All Point Clouds

Python 712 52 Updated Jul 1, 2026

REFLEX Dataset: A Multimodal Dataset of Human Reactions to Robot Failures and Explanations

Python 6 Updated Mar 2, 2025

Official code for CVPR 2026 paper: VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection

Python 145 4 Updated Jul 15, 2026
Next