Skip to content
View wjpoom's full-sized avatar

Block or report wjpoom

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Official implementation of: Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

Python 11 Updated Jun 11, 2026

ARM: An AutoRegressive Large Multimodal Model with Discrete Representations

50 Updated Jun 10, 2026

Code for RepWAM: World Action Modeling with Representation Visual-Action Tokenizers

59 1 Updated Jun 14, 2026

[CVPR 2026] FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding

Python 73 1 Updated Mar 16, 2026
Python 6 1 Updated Jun 9, 2026

Decoupled Memory Selection for Multi-target Video Segmentation of SAM3

Python 58 4 Updated Jan 16, 2026

[ICML 2026] The official implementation of paper "Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key to Unification"

Python 46 Updated Jul 13, 2026

This repository contains the implementation of SAM3 trackers.

Python 36 3 Updated Jun 30, 2026

The official repository of Qwen-VLA

718 25 Updated May 29, 2026

[CVPR-26] Official repository of "Compositional Text-to-Image Generation via Region-Aware Bimodal Direct Preference Optimization“

Python 2 Updated May 29, 2026

A curated, continuously updated reading list, paper blogs, and resources for World Action Models (WAMs) in embodied AI.

HTML 1,168 30 Updated Jul 23, 2026

[CVPR-26] Official repository of "CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization"

Python 19 Updated Mar 9, 2026

Repo for Qwen Image Finetune

Jupyter Notebook 249 27 Updated Jun 8, 2026
Python 5 Updated Jan 23, 2026

Open-source red teaming framework for MLLMs with 42+ attack methods

Python 258 20 Updated Jul 17, 2026

StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

Python 3,279 417 Updated Jul 20, 2026

[ICLR'26] Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?

Python 52 1 Updated Mar 9, 2026

Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.

1,493 47 Updated Mar 9, 2026

Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better

Python 190 18 Updated Apr 7, 2026

📖 This is a repository for organizing papers, codes, and other resources related to unified multimodal models.

365 14 Updated Jan 8, 2026

[ACL 2025] "World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning." https://arxiv.org/abs/2503.10480

Jupyter Notebook 18 Updated Jul 22, 2025

Pytorch implementation for the paper titled "SimpleAR: Pushing the Frontier of Autoregressive Visual Generation"

Python 431 25 Updated Jun 20, 2025

Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

Python 202 19 Updated Apr 9, 2026

This is a curated list of "Embodied AI or robot with Large Language Models" research. Watch this repository for the latest updates! 🔥

1,837 98 Updated Jul 14, 2026

EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL

Python 5,081 383 Updated Jul 23, 2026

This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!

1,435 64 Updated May 11, 2026

[WAICA-26 Best Student Paper] Official repository of "Enhancing Vision Foundation Models via Multimodal Continual Pre-Training"

Python 49 2 Updated Jul 21, 2026

[TPAMI 2025] Towards Visual Grounding: A Survey

Shell 322 25 Updated Nov 18, 2025

[ICCV 2025] MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance

Python 184 13 Updated Feb 11, 2026

A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.

HTML 679 42 Updated Jul 22, 2026
Next