Stars
[ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.
[CVPR 2026] Official codes of "Monet: Reasoning in Latent Visual Space Beyond Image and Language"
We propose an efficient flow-based multimodal generation model with bidirectional flows.
Open-source native multimodal pretraining — without catastrophic forgetting.
SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
[ICLR 2026] This is an early exploration to introduce Interleaving Reasoning to Text-to-image Generation field and achieve the SoTA benchmark performance. It also significantly improves the quality…
[CVPR 2026] Official repo for "EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation"
LLaDA2.0-Uni: Understanding and Generation the World.
MonkeyOCRv2 Vision Encoder — A Document-Native Visual Backbone
Official implementation of DeltaV, a unified multimodal model for interleaved reasoning with visual state updates.
(NeurIPS 2025) Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
Reverse Chain-of-Thought Problem Generation for Geometric Reasoning in Large Multimodal Models
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
Spatial Aptitude Training for Multimodal Langauge Models
MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
Multimodal OCR: Parse Anything from Documents
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
Official electron build of draw.io
Towards Efficient Multimodal Large Language Models: A Survey on Token Compression
[ICLR26] ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
[ICLR 2026] The official repository for paper "ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning"
Deformable ConvNets V2 (DCNv2) in PyTorch
[ICCV 2023] Official implementation of the paper "DFA3D: 3D Deformable Attention For 2D-to-3D Feature Lifting"
Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence