UCP Checker
UCP Playground

Agent shopping sessions, observed.

UCP Playground runs real AI models against real UCP-verified stores and logs every JSON-RPC tool call. The aggregate across 26 models from 8 vendors and 531 stores tells us what works in agentic shopping — and what doesn't.

Snapshot: Sep 29, 2026
26
Models Tested
Frontier + open-weight
8
Model Vendors
Anthropic, Google, OpenAI, Meta, DeepSeek, xAI, Alibaba, Moonshot
531
Stores Tested
Shopify, Woo, Magento, custom
25.3%
Reach Checkout
Of all agent sessions
What it is
01
An agent runtime, not a benchmark suite

Pick a model. Pick a UCP-verified store. The agent connects to the store's MCP endpoint, browses, carts, and tries to check out — using only the tools the store advertises in its /.well-known/ucp profile.

02
Every tool call is logged and replayable

Every JSON-RPC request and response, every retry, every failed schema match — captured per session. Postman for agentic commerce: inspect what an agent actually did, not what it claimed to do.

03
Independent of any vendor

Models from Anthropic, Google, OpenAI, Meta, DeepSeek, xAI, Alibaba, Moonshot run side-by-side against the same stores. Vendors can't credibly benchmark themselves; the platform layer has the same problem one level down.

What the sessions reveal
14.1%
Gemini 3.6 Flash leads the leaderboard

Roughly 1 in 7 sessions. Fast, decisive tool-call rhythm — not deliberation — is what wins at agentic shopping.

↓
Reasoning-tuned models underperform

DeepSeek R1, QwQ, Grok 3 Mini consistently burn tokens on chain-of-thought and miss the next tool call. Agentic shopping rewards speed, not introspection.

75/25
Store implementation drives ~75% of variance

A well-typed profile and tight tool responses move the conversion needle more than swapping models. Most "the agent is broken" tickets are actually store-side.

Session funnel
Checkout reached
25.3%
Cart created
21.5%
Search only
29.7%
Failed
23.2%
Model leaderboard
Model Avg tokens Avg duration Session share
Gemini 3.6 Flash 74,954 22.2s 14.1%
Gemini 3 Flash 76,507 22.5s 13.2%
Claude Sonnet 4.5 69,836 33.8s 12%
Claude Opus 4.6 61,037 30.5s 7.9%
Gemini 2.5 Flash 44,976 13.9s 6.8%
GPT-5.2 52,119 33.2s 6.7%
Gemini 3.1 Pro 61,504 42.4s 5.3%
GPT-4o 34,077 16.4s 5.1%
Gemini 2.5 Pro 47,160 37.4s 5.1%
DeepSeek v3.2 68,697 47.2s 3.1%
o4-mini 70,285 39.6s 2.6%
Grok 4 35,601 45.4s 2.6%
Claude Sonnet 5 261,470 47.2s 2.5%
Llama 3.3 70B 58,192 40.4s 2.4%
DeepSeek R1 23,023 53.7s 1.7%
DeepSeek v4 115,018 67.1s 1.6%
Claude Opus 5 194,451 54.1s 1.6%
Grok 3 Mini 35,632 40.6s 1%
QwQ 32B 14,586 36.7s 0.9%
Llama 4 Maverick 43,535 15.1s 0.9%
DeepSeek v4 Flash 124,257 49.9s 0.7%
Qwen3.7 Max 72,412 27.4s 0.7%
Grok 4.5 110,646 41.1s 0.7%
Kimi K3 77,433 84.3s 0.6%
Gemini 3.8 Flash 120,674 60.7s 0.2%
GPT-6 Luna 103,889 21s 0.1%
Tool call patterns
Tool Call share Avg latency Error rate
search_catalog 35.1% 725ms 21.9%
update_cart 20.1% 609ms 32.5%
search_shop_catalog 15.5% 483ms 26.9%
search_global_products 7% 347ms 0%
get_product_details 5.8% 266ms 25.1%
get_cart 3.9% 279ms 43%
create_cart 3.4% 916ms 36%
POST /catalog/search 3.4% 4,543ms 37.4%
search_products 3.1% 1,229ms 14.3%
get_product 2.7% 1,043ms 35.4%
FAQ

Common questions

What is UCP Playground?
UCP Playground is an agent shopping runtime at ucpplayground.com. It runs real AI models against real UCP-verified stores, logs every JSON-RPC tool call, and makes every session replayable. Think of it as Postman for agentic commerce — observability for agent behaviour, not a vendor benchmark.
How is UCP Playground different from UCP Checker?
UCP Checker evaluates your profile, surface signals, and lightweight probes — fast, safe, runs at scale. UCP Playground runs full end-to-end agent shopping simulations against your live infrastructure. Checker tells you "can agents find and parse you" — Playground tells you "can agents complete a transaction".
Is UCP Playground a UCP demo?
Yes — it functions as a live UCP demo. Every session is a real AI model exercising real UCP capabilities (search, cart, checkout) against a real store, in real time. Unlike a recorded demo video, every Playground session is reproducible: pick the same model and store and run it again. Several stores publish demo profiles specifically for this purpose (demo-travel.ucp.dev, sandbox subdomains) so you can watch UCP in action without affecting production inventory.
Which AI model performs best at agentic shopping?
Gemini 3.6 Flash leads the leaderboard with 14.1% of total session share. The pattern across 26 models from 8 vendors is consistent: agentic shopping rewards fast, decisive tool-use rhythm, not deliberation.
Why do reasoning-tuned models underperform on shopping tasks?
Reasoning-tuned models (DeepSeek R1, QwQ, Grok 3 Mini, o4-mini) burn tokens on chain-of-thought before each tool call. Shopping is a fast tool-use rhythm — search → details → cart → checkout — and the deliberation overhead causes them to drop the thread, miss the next call, or hit schema mismatches. Frontier non-reasoning models outperform on the same stores.
How is the checkout completion rate calculated?
Sessions that reach a checkout state divided by total sessions. Currently 25.3% of sessions reach checkout. The remainder split between cart-created, search-only, and failed sessions — the funnel above shows the full breakdown.
How many stores has UCP Playground tested?
531 unique UCP-verified stores as of 2026-09-29. Sessions span Shopify, WooCommerce, Magento, custom Rails apps, and bespoke MCP servers. The same agent code runs against all of them — anything different in the outcome is the store, not the agent.
Does my store implementation matter more than the model I pick?
Yes — by a wide margin. Across the dataset, store implementation drives roughly 75% of variance in agent outcomes; model choice drives the remaining ~25%. A well-typed profile, fast and consistent tool responses, and complete capability coverage move the conversion needle more than swapping to a "better" model. Most "the agent is broken" tickets are actually store-side.
Is UCP Playground free to use?
Yes — running sessions is free at ucpplayground.com. Pick a model, pick a store, watch the agent shop. There is also a headless API for running sessions in CI as part of the paid tier.
How do I run a UCP test on my own store?
To run a UCP test against your own store, point Playground at your domain. If you have a valid /.well-known/ucp profile and reachable MCP or REST endpoints, agent sessions will run against it directly. Stores that grade well on UCP Score typically also pass UCP tests in Playground, but a Playground test surfaces runtime issues that static checks cannot — schema drift, slow tool responses, partial cart state, missing variant data.
UCP Playground

Run your own agent session

Pick a model, pick a store, watch the agent shop. Every JSON-RPC tool call is logged and every session is replayable — so you can see what an agent actually did, not just what it claimed.

Open UCP Playground
Free to use 26 models available Real verified stores

Building your own agent? Start at /agents · Reading the spec? /protocol

Agent Session
claude-sonnet-4-5 · everlane.com · 38s
Checkout reached
search_shop_catalog
update_cart
create_checkout
Weekly UCP Report

Get the agentic commerce digest every Monday

Real adoption data, ecosystem trends, new spec versions, and the stores that broke or recovered this week. Read by founders and engineers building the next generation of commerce.

Free forever No spam Unsubscribe anytime

View a sample report →

Weekly UCP Report
Issue #53 · Oct 5, 2026
+664
new verified stores
Verified rate
Latest spec
Cart capability