Stars
[CVPR 2026] Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
[ACL'26 Findings] OmniDiagram: Advancing Unified Diagram Code Generation via Visual Interrogation Reward
slime is an LLM post-training framework for RL Scaling.
World Models for Policy Refinement in StarCraft II
Code for 🌍 UI-Simulator: LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
🌎💪 BrowserGym, a Gym environment for web task automation
[AAAI'26] Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks
Official Implementation for the paper "VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models"
[CVPR 2026] Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
SynthAgent: Adapting Web Agents with Synthetic Supervision, ACL 2026
[ICML'24] SeeAct is a system for generalist web agents that autonomously carry out tasks on any given website, with a focus on large multimodal models (LMMs) such as GPT-4V(ision).
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
[NeurIPS'25 Spotlight🔥] Official Implementation of RobustMerge: Parameter-Efficient Model Merging for MLLMs with Direction Robustness
Tongyi Deep Research, the Leading Open-source Deep Research Agent
GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's TerminalBench leaderboard.
[ACL 2025 Oral] The official repository of our paper: CADReview: Automatically Reviewing CAD Programs with Error Detection and Correction
[ICLR 2026] Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner