Stars
DeepSeek Harness: Everything is a Plugin.
Gemma-based Multilingual Machine Translation Models
A high-throughput and memory-efficient inference and serving engine for LLMs
[ECCV'26] Official implementation of paper "GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models".
[Technical Report] An End-to-End Multimodal GUI Agent for Real Mobile Environments
[2026-ECCV] UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation
Native macOS control center for AI coding agents — monitor sessions, approve actions, and jump back instantly.
Make raster image ediatble. 将生成图像(如GPT-Image-2或Nano Banana)或截图变为可编辑的格式,实现高质量PPT、论文图直出。
[ICML 2026] ZwZ model family: SOTA fine-grained perception performace; ZoomBench: a new challenging perception benchmark
Digital Agents Meet World Models: A Survey
Reference PyTorch implementation and models for DINOv3
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
My learning notes for ML SYS.
Post-training with Tinker
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
CDLA: A Chinese document layout analysis (CDLA) dataset
[NeurIPS 2025] Implementation of the paper "BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent"
DART-GUI: Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
Mobile-Agent: The Powerful GUI Agent Family
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
Thinking with Videos from Open-Source Priors. We reproduce chain-of-frames visual reasoning by fine-tuning open-source video models. Give it a star 🌟 if you find it useful.
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
Official implementation of TrajBooster
Official code repository of Shuffle-R1