-
Fudan University
- Fudan University, Shanghai, China
- https://wjpoom.github.io/
Lists (5)
Sort Name ascending (A-Z)
Stars
Official implementation of: Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback
ARM: An AutoRegressive Large Multimodal Model with Discrete Representations
Code for RepWAM: World Action Modeling with Representation Visual-Action Tokenizers
[CVPR 2026] FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
Decoupled Memory Selection for Multi-target Video Segmentation of SAM3
[ICML 2026] The official implementation of paper "Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key to Unification"
This repository contains the implementation of SAM3 trackers.
[CVPR-26] Official repository of "Compositional Text-to-Image Generation via Region-Aware Bimodal Direct Preference Optimization“
A curated, continuously updated reading list, paper blogs, and resources for World Action Models (WAMs) in embodied AI.
[CVPR-26] Official repository of "CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization"
Repo for Qwen Image Finetune
Open-source red teaming framework for MLLMs with 42+ attack methods
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
[ICLR'26] Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
📖 This is a repository for organizing papers, codes, and other resources related to unified multimodal models.
[ACL 2025] "World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning." https://arxiv.org/abs/2503.10480
Pytorch implementation for the paper titled "SimpleAR: Pushing the Frontier of Autoregressive Visual Generation"
Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
This is a curated list of "Embodied AI or robot with Large Language Models" research. Watch this repository for the latest updates! 🔥
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!
[WAICA-26 Best Student Paper] Official repository of "Enhancing Vision Foundation Models via Multimodal Continual Pre-Training"
[TPAMI 2025] Towards Visual Grounding: A Survey
[ICCV 2025] MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.