Ratel Benchmark

Ratel, measured

No corporate fluff. Just raw data and sharp claws.

Evaluated with golden industry standards, measured with BFCL v3.

Token Spend by Model

Mean total tokens per task, one row per model. Oracle = only the gold tools in context (the floor); With Ratel = the search_tools gateway; Baseline = the full tool pool. The dumbbell shows how far Ratel pulls each model back toward the oracle floor.

010.0k20.0k30.0kClaude Sonnet 4.629.3k1.8k3.7kQwen3 4b28.1k7142.5kClaude Haiku 4.525.6k1.4k3.7kGPT 5.4 Mini15.0k4311.7k
Oraclegold tools onlyWith Ratelsearch_tools gatewayBaselinefull tool pool

How Ratel effect agent accuracy and token cost

Pick a model. The bars compare mean total tokens per task; the radar overlays all five metrics for With Ratel (the search_tools gateway) vs Without Ratel (the full tool pool). On Token efficiency the axis shows tokens spent from the center out, so Without Ratel spikes there. Hover any plot for exact values.

Mean total tokens per task

87% tokens
Without RatelClaude Sonnet 4.6 · full tool pool
0
With RatelClaude Sonnet 4.6 · search_tools gateway
0
Anthropic

Claude Sonnet 4.6

Task completionTool selectionRecallToken efficiencySpeed
Task completion
92.8%
Without Ratel 92.0%
Tool selection
97.2%
Without Ratel 98.0%
Recall
95.1%
Without Ratel 95.5%
Token cost
3,708
Without Ratel 29,301
−87%
p50 latency
7,900 ms
Without Ratel 8,707 ms

Tool Selection by LLM

Pooled over the simple + multiple splits. Without Ratel = full tool pool in context; With Ratel = search_tools gateway; Oracle = only the gold tools in context (the ceiling).

ModelArmTask completionTool selectionRecallMean total tokens
Claude Haiku 4.5Without Ratel79.3%84.1%82.2%25,606
Oracle75.5%80.1%78.3%1,386
With Ratel91.5%96.2%94.1%3,653
Claude Sonnet 4.6Without Ratel92.0%98.0%95.5%29,301
Oracle94.3%99.3%96.9%1,829
With Ratel92.8%97.2%95.1%3,708
GPT 5.4 MiniWithout Ratel84.8%96.5%93.0%15,006
Oracle88.3%98.8%95.9%431
With Ratel84.3%95.3%92.1%1,671
Qwen3 4bWithout Ratel90.2%95.8%93.6%28,060
Oracle95.3%99.3%97.8%714
With Ratel86.0%91.7%89.1%2,540

Retrieval evaluation

Evaluation of retriever engine. Pick a pool size — the bars and tables show the mean ranking metrics at top k = 1,3,5. Accuracy = a gold tool in the top-K; complete = every gold tool in the top-K; gold = share of queries whose gold tool was retrievable.

Tool pool size

Simple (single gold tool)

100%98%96%94%92%90%
Accuracy
Complete
Recall
MRR
nDCG
K=1
Accuracy
Complete
Recall
MRR
nDCG
K=3
Accuracy
Complete
Recall
MRR
nDCG
K=5
KaccuracycompleterecallMRRnDCGgold sim.
197.0%97.0%0.9700.9700.97097.0%
399.0%99.0%0.9900.9790.98299.0%
599.5%99.5%0.9950.9800.98499.5%

Multiple (several gold tools)

100%98%96%94%92%90%
Accuracy
Complete
Recall
MRR
nDCG
K=1
Accuracy
Complete
Recall
MRR
nDCG
K=3
Accuracy
Complete
Recall
MRR
nDCG
K=5
KaccuracycompleterecallMRRnDCGgold sim.
196.5%96.5%0.9650.9650.96596.5%
3100.0%100.0%1.0000.9810.986100.0%
5100.0%100.0%1.0000.9810.986100.0%