Skip to content
View Pengjie-W's full-sized avatar

Block or report Pengjie-W

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

[ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.

Python 380 33 Updated Mar 29, 2026

A unified multimodal model toolkit

Python 679 135 Updated Aug 12, 2026

[CVPR 2026] Official codes of "Monet: Reasoning in Latent Visual Space Beyond Image and Language"

Python 216 10 Updated Mar 19, 2026

We propose an efficient flow-based multimodal generation model with bidirectional flows.

Jupyter Notebook 18 Updated Feb 18, 2026

Open-source native multimodal pretraining — without catastrophic forgetting.

Python 15 1 Updated Jul 2, 2026

SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles

Python 4,820 410 Updated Aug 15, 2026

[ICLR 2026] This is an early exploration to introduce Interleaving Reasoning to Text-to-image Generation field and achieve the SoTA benchmark performance. It also significantly improves the quality…

Python 101 1 Updated Jan 26, 2026

[CVPR 2026] Official repo for "EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation"

Python 65 2 Updated Mar 13, 2026
Python 30 1 Updated Jul 25, 2026

LLaDA2.0-Uni: Understanding and Generation the World.

Python 724 46 Updated May 29, 2026

MonkeyOCRv2 Vision Encoder — A Document-Native Visual Backbone

Python 875 83 Updated Aug 13, 2026

Official implementation of DeltaV, a unified multimodal model for interleaved reasoning with visual state updates.

Python 32 Updated Jul 13, 2026

🔨AI 方向好用的科研工具

3,415 405 Updated Jun 10, 2024

(NeurIPS 2025) Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

Python 77 Updated May 21, 2026

Reverse Chain-of-Thought Problem Generation for Geometric Reasoning in Large Multimodal Models

Python 213 8 Updated Nov 4, 2024

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

Python 4,340 744 Updated Aug 14, 2026

Spatial Aptitude Training for Multimodal Langauge Models

Python 33 1 Updated Feb 8, 2026

MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling

Python 10 Updated Apr 20, 2026

Multimodal OCR: Parse Anything from Documents

Python 321 24 Updated Mar 20, 2026

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 195,050 109,115 Updated Aug 6, 2026

Official electron build of draw.io

JavaScript 62,629 5,747 Updated Aug 14, 2026

Open-source unified multimodal model

Python 6,146 547 Updated May 4, 2026

Towards Efficient Multimodal Large Language Models: A Survey on Token Compression

219 9 Updated Aug 10, 2026

[ICLR26] ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding

Python 89 1 Updated Mar 20, 2026

[ICLR 2026] The official repository for paper "ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning"

Jupyter Notebook 192 17 Updated May 1, 2026

OCR in the Era of Large Language Models

696 50 Updated Aug 14, 2026

Deformable ConvNets V2 (DCNv2) in PyTorch

Cuda 1,484 230 Updated Nov 18, 2022

[ICCV 2023] Official implementation of the paper "DFA3D: 3D Deformable Attention For 2D-to-3D Feature Lifting"

Python 184 4 Updated Apr 12, 2025

Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

Python 395 20 Updated Jun 20, 2026
Next