AI Lab · Track 3

Eval Dashboard

Run a benchmark

Fires every test question against your live RAG or agent API, then scores accuracy, latency, and cost.

Loading suites…

Results

Run a suite to see per-question scores.