Lists (3)
Sort Name ascending (A-Z)
Stars
Official implementation of DeltaV, a unified multimodal model for interleaved reasoning with visual state updates.
MonkeyOCRv2 Vision Encoder — A Document-Native Visual Backbone
Reverse Chain-of-Thought Problem Generation for Geometric Reasoning in Large Multimodal Models
MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
🚀 2026届大模型算法岗实习面经 | 包含 DeepSeek/Qwen 技术报告解析、手撕 PPO/RoPE/Transformer、RLHF 核心与八股文 | 持续更新中...
[ICLR26] ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Official repository for the UAE paper, unified-GRPO, and unified-Bench
A lightweight LMM-based Document Parsing Model
[ICLR 2026] OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
Official Repository of "Learning to Reason under Off-Policy Guidance"
Official code implementation of Slow Perception:Let's Perceive Geometric Figures Step-by-step
Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 Highlight)