Skip to content
View dle666's full-sized avatar

Block or report dle666

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Official implementation of DeltaV, a unified multimodal model for interleaved reasoning with visual state updates.

Python 32 Updated Jul 13, 2026

MonkeyOCRv2 Vision Encoder — A Document-Native Visual Backbone

Python 853 82 Updated Aug 13, 2026

OCR in the Era of Large Language Models

695 50 Updated Aug 14, 2026

Reverse Chain-of-Thought Problem Generation for Geometric Reasoning in Large Multimodal Models

Python 213 8 Updated Nov 4, 2024
Python 27 Updated Jul 5, 2026

MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling

Python 10 Updated Apr 20, 2026

🚀 2026届大模型算法岗实习面经 | 包含 DeepSeek/Qwen 技术报告解析、手撕 PPO/RoPE/Transformer、RLHF 核心与八股文 | 持续更新中...

635 12 Updated Mar 28, 2026

[ECCV 26] Video Streaming Thinking

Python 120 3 Updated Jul 28, 2026

[ICLR26] ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding

Python 89 1 Updated Mar 20, 2026

Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights

JavaScript 32 1 Updated Jan 9, 2026

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

Python 1,920 150 Updated Jul 29, 2026

Xiaomi Miloco

Python 3,235 285 Updated Aug 14, 2026

Official repository for the UAE paper, unified-GRPO, and unified-Bench

Python 166 7 Updated Sep 12, 2025

[ICCV 2025] LIRA

Python 22 4 Updated Nov 25, 2025

A lightweight LMM-based Document Parsing Model

Python 6,625 459 Updated Jul 20, 2026

[ICLR 2026] OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

Python 77 3 Updated May 26, 2026

Official Repository of "Learning to Reason under Off-Policy Guidance"

Python 463 72 Updated Mar 20, 2026

Official code implementation of Slow Perception:Let's Perceive Geometric Figures Step-by-step

Python 162 8 Updated Jul 28, 2025

Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 Highlight)

Python 1,951 139 Updated Jun 2, 2026