Claude Sonnet 4.6
- Task completion
- 92.8%
- Without Ratel 92.0%
- Tool selection
- 97.2%
- Without Ratel 98.0%
- Recall
- 95.1%
- Without Ratel 95.5%
- Token cost
- 3,708
- Without Ratel 29,301
- −87%
- p50 latency
- 7,900 ms
- Without Ratel 8,707 ms
No corporate fluff. Just raw data and sharp claws.
Evaluated with golden industry standards, measured with BFCL v3.
Mean total tokens per task, one row per model. Oracle = only the gold tools in context (the floor); With Ratel = the search_tools gateway; Baseline = the full tool pool. The dumbbell shows how far Ratel pulls each model back toward the oracle floor.
Pick a model. The bars compare mean total tokens per task; the radar overlays all five metrics for With Ratel (the search_tools gateway) vs Without Ratel (the full tool pool). On Token efficiency the axis shows tokens spent from the center out, so Without Ratel spikes there. Hover any plot for exact values.
Pooled over the simple + multiple splits. Without Ratel = full tool pool in context; With Ratel = search_tools gateway; Oracle = only the gold tools in context (the ceiling).
| Model | Arm | Task completion | Tool selection | Recall | Mean total tokens |
|---|---|---|---|---|---|
| Claude Haiku 4.5 | Without Ratel | 79.3% | 84.1% | 82.2% | 25,606 |
| Oracle | 75.5% | 80.1% | 78.3% | 1,386 | |
| With Ratel | 91.5% | 96.2% | 94.1% | 3,653 | |
| Claude Sonnet 4.6 | Without Ratel | 92.0% | 98.0% | 95.5% | 29,301 |
| Oracle | 94.3% | 99.3% | 96.9% | 1,829 | |
| With Ratel | 92.8% | 97.2% | 95.1% | 3,708 | |
| GPT 5.4 Mini | Without Ratel | 84.8% | 96.5% | 93.0% | 15,006 |
| Oracle | 88.3% | 98.8% | 95.9% | 431 | |
| With Ratel | 84.3% | 95.3% | 92.1% | 1,671 | |
| Qwen3 4b | Without Ratel | 90.2% | 95.8% | 93.6% | 28,060 |
| Oracle | 95.3% | 99.3% | 97.8% | 714 | |
| With Ratel | 86.0% | 91.7% | 89.1% | 2,540 |
Evaluation of retriever engine. Pick a pool size — the bars and tables show the mean ranking metrics at top k = 1,3,5. Accuracy = a gold tool in the top-K; complete = every gold tool in the top-K; gold = share of queries whose gold tool was retrievable.
| K | accuracy | complete | recall | MRR | nDCG | gold sim. |
|---|---|---|---|---|---|---|
| 1 | 97.0% | 97.0% | 0.970 | 0.970 | 0.970 | 97.0% |
| 3 | 99.0% | 99.0% | 0.990 | 0.979 | 0.982 | 99.0% |
| 5 | 99.5% | 99.5% | 0.995 | 0.980 | 0.984 | 99.5% |
| K | accuracy | complete | recall | MRR | nDCG | gold sim. |
|---|---|---|---|---|---|---|
| 1 | 96.5% | 96.5% | 0.965 | 0.965 | 0.965 | 96.5% |
| 3 | 100.0% | 100.0% | 1.000 | 0.981 | 0.986 | 100.0% |
| 5 | 100.0% | 100.0% | 1.000 | 0.981 | 0.986 | 100.0% |