Skip to content
View QinYang79's full-sized avatar

Block or report QinYang79

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

[ACM MM 2026] Official resources of "Multi-Branch Policy Optimization for Multimodal Large Language Models".

Python 2 1 Updated Aug 11, 2026

Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.

1,504 47 Updated Mar 9, 2026

⭐️ A cross-platform CLI All-in-One assistant tool for Claude Code, Codex & Gemini CLI.

Rust 4,702 275 Updated Aug 13, 2026

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

Rust 127,691 8,734 Updated Aug 17, 2026

[CVPR 2026 Highlight] DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding

Python 19 1 Updated Jun 4, 2026

ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…

Python 14,786 1,300 Updated Aug 15, 2026

Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps

Python 47,076 8,318 Updated Aug 17, 2026

100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.

Python 132,916 19,557 Updated Aug 17, 2026

[NeurIPS 2025] First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training

Python 88 2 Updated Oct 29, 2025
Python 53 6 Updated Oct 10, 2025

Tongyi Deep Research, the Leading Open-source Deep Research Agent

Python 19,833 1,506 Updated Feb 27, 2026

The paper list of the 86-page SCIS cover paper "The Rise and Potential of Large Language Model Based Agents: A Survey" by Zhiheng Xi et al.

8,172 495 Updated Sep 12, 2025

Deep Research

Python 303 10 Updated Aug 26, 2025

UI-Venus is a native UI agent designed to perform precise GUI element grounding and effective navigation using only screenshots as input.

Python 1,007 83 Updated May 11, 2026

[ICLR2026] This is the first paper to explore how to effectively use R1-like RL for MLLMs and introduce Vision-R1, a reasoning MLLM that leverages cold-start initialization and RL training to incen…

Python 1,570 27 Updated Mar 20, 2026

This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!

1,440 69 Updated Aug 2, 2026

[AAAI 2026] GUI-G²: Gaussian Reward Modeling for GUI Grounding

Python 311 10 Updated Apr 15, 2026

A library for advanced large language model reasoning

Python 2,341 203 Updated Jun 10, 2025

EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL

Python 5,117 384 Updated Jul 30, 2026

Trustworthy Visual-Textual Retrieval (TIP 2025 Pytorch Code)

Python 9 Updated Jan 14, 2026

[CVPR 2024 Highlight] Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding

Python 412 27 Updated Oct 7, 2024

[CVPR 2024 Highlight] OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation

Python 415 31 Updated Aug 24, 2024

[CVPR 2025] Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens

Python 84 12 Updated Oct 9, 2025

[ICML 2025] Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

Python 14 Updated May 28, 2025

📖 A curated list of resources dedicated to hallucination of multimodal large language models (MLLM).

1,037 47 Updated Sep 27, 2025

😎 A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, Agent, and Beyond

357 14 Updated Jan 22, 2026

The code will come soon.

Python 11 3 Updated Aug 27, 2025

GEA: Generation-Enhanced Alignment for Text-to-Image Person Retrieval

6 Updated Jul 14, 2025

Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.

Python 27,511 2,038 Updated Jan 9, 2026
Next