Skip to content
View yyyybq's full-sized avatar

Block or report yyyybq

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects

Jupyter Notebook 302 15 Updated Jun 11, 2026

Holistic Evaluation of Multimodal LLMs on Spatial Intelligence

Python 118 10 Updated Jul 1, 2026

A Curated List of Awesome Works in World Modeling, Aiming to Serve as a One-stop Resource for Researchers, Practitioners, and Enthusiasts Interested in World Modeling.

3,219 137 Updated Jul 24, 2026

This is a collection of recent papers on reasoning in video generation models.

165 6 Updated Jul 21, 2026

Public code for XFactor: Introduces the first geometry-free model to achieve true self-supervised / pose-free Novel View Synthesis (NVS) by learning transferable latent camera pose representations.

Python 160 3 Updated May 11, 2026

One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

Python 4,335 626 Updated Jul 23, 2026

A collection of papers on semantic correspondence, organized by year.

32 2 Updated Dec 10, 2025

[ICLR'26] IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction

Python 427 25 Updated Dec 1, 2025

[Awesome-Spatial-VLMs] This repository is the official, community-maintained resource for the survey paper: Spatial Intelligence in Vision-Language Models: A Comprehensive Survey;

Python 526 3 Updated Apr 10, 2026

The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.

Python 5,597 568 Updated Jul 24, 2026

Official implementation of paper "Unified World Models: Memory-Augmented Planning and Foresight for Visual Navigation"

Python 300 34 Updated May 16, 2026

[ECCV 2026🔥] SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

Python 93 7 Updated Nov 26, 2025

Open-source unified multimodal model

Python 6,117 544 Updated May 4, 2026

https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT

Python 137 7 Updated Jan 30, 2026

Official Repository for “CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models" [CVPR2025]

Python 5 Updated Dec 14, 2025

InteriorGS: 3D Gaussian Splatting Dataset of Semantically Labeled Indoor Scenes

283 10 Updated Aug 4, 2025

RAGEN leverages reinforcement learning to train LLM reasoning agents in interactive, stochastic environments.

Python 2,756 229 Updated Jul 24, 2026

World model reasoning RL for multi-turn VLM agents

Python 488 61 Updated Jul 23, 2026
Python 163 6 Updated Mar 23, 2026

The first collection of academic iKUN papers in the world

5 1 Updated Jun 21, 2024

Github repository for "Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas" (ICML 2025)

Python 76 11 Updated May 2, 2025

Official code for Paper "Mantis: Multi-Image Instruction Tuning" [TMLR 2024 Best Paper]

Python 239 23 Updated Jan 3, 2026

A paper list for spatial reasoning

766 44 Updated Jan 19, 2026

本人的科研经验

13,407 673 Updated Jun 6, 2026

Kolmogorov Arnold Networks

Jupyter Notebook 16,327 1,565 Updated Jan 19, 2025
Python 9 1 Updated Apr 30, 2024

TheaterGen: Character Management with LLM for Consistent Multi-turn Image Generation

Python 69 9 Updated Sep 26, 2024
Next