TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
-
Updated
Sep 8, 2026 - C++
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
ComfyUI-QwenVL custom node: Integrates the Qwen-VL series, including Qwen2.5-VL, Qwen3-VL, Qwen3.5-VL, Qwen3.6-VL (MoE), and Qwen3.8-VL, with GGUF support for advanced multimodal AI in text generation, image understanding, and video analysis.
A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.
A minimal codebase for finetuning large multimodal models, supporting llava-1.5/1.6, llava-interleave, llava-next-video, llava-onevision, llama-3.2-vision, qwen-vl, qwen2-vl, phi3-v etc.
Reinforcement Learning of Vision Language Models with Self Visual Perception Reward
给 DeepSeek 装上眼睛 — MCP Server + 通义千问VL, 剪贴板图片→视觉模型→文字描述 / Give DeepSeek the ability to see images via clipboard + Qwen-VL
Mark web pages for use with vision-language models
Self-evolving agentic reward framework for image-editing evaluation — 47.4% on EditReward-Bench from only 100 preference demos, no reward-model training. arXiv 2605.08703.
Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches the mock. Ships the benchmark behind that judgement — four fixtures, 23 injected defects, and every raw model transcript. Retires itself when DeepSeek ships vision.
Local Video RAG Engine. A FastAPI microservice for video understanding: Scene Detection + Whisper ASR + Qwen3-VL. Optimized for Apple Silicon (MLX) & Windows/Linux (Llama.cpp).
An AI Agent that is able to control your screen to complste any task
Give non-multimodal Claude Code main models the ability to see pasted screenshots — a ~200-line UserPromptSubmit hook.
🎬 Extract AI prompts from video using Vision LLM (llama.cpp API) — Gradio WebUI + CLI
基于Qwen2.5-VL-3B + QLoRA的智能驾驶场景结构化理解系统
PriorTR (ECCV 2026): training-free, prior-corrected visual token reduction for accelerating multimodal LLMs — image & video.
Model-agnostic AI poster grid inspector: detect a poster's layout grid system with a vision model, view it overlaid, and export SVG guides / CSS Grid / Tailwind code. Ships as an installable agent skill.
Vision-language gateway plugin for DeepSeek Harness - paste an image, DeepSeek sees text
基于 Qwen3-VL-Flash 视觉语言模型的 Web 端自动化图像标注平台,支持 2D 物体检测、3D 空间定位(9-DOF)与视角遮挡分析,可导出 COCO / Pascal VOC 格式。
A robotic sequential grasping system integrating YOLO detection and Qwen-VLM fine-tuning, enabling a full loop from manual teaching to LLM-based logical manipulation.
To associate your repository with the qwen-vl topic, visit your repo's landing page and select "manage topics."