Skip to content

[evals] AD2 full-profile on-demand discrimination (follow-up) - #13

Draft
patrykkopycinski wants to merge 3 commits into
feat/attack-discovery-agent-builder-evalsfrom
feat/ad2-full-profile-evals
Draft

patrykkopycinski wants to merge 3 commits into
feat/attack-discovery-agent-builder-evalsfrom
feat/ad2-full-profile-evals

Conversation

@patrykkopycinski

@patrykkopycinski patrykkopycinski commented Jul 14, 2026 •

Copy link
Copy Markdown
Owner

Summary

Follow-up to elastic/kibana#277625 — adds the full profile portable-seeder cohort as an on-demand eval layer with explicit FPR/discrimination gates, and restores attack-discovery-generator skill hardening moved out of elastic#277625 so baseline Δ is measurable.

Does not change weekly CI (golden-path EVAL_GREP unchanged).

Stacking: This PR targets feat/attack-discovery-agent-builder-evals on the fork. After elastic#277625 merges, retarget to main on elastic/kibana (or rebase and open a new PR there).

Eval profiles in this change

Profile Spec Cadence Purpose
Golden-path attack_discovery_agent_builder.spec.ts Weekly (EVAL_GREP) Routing + workflow plumbing (elastic#277625)
Clean profile clean_profile_provided_alerts.spec.ts On-demand Quality on 4 realistic chains (elastic#277625)
Full profile full_profile_discrimination.spec.ts On-demand only Signal vs noise under ~150 distractor alerts

What changed

  • Scenario registry full profile: 7 signal chains (clean + AWS/Azure/macOS) + deterministic 110 background + 40 Defender loud-cluster alerts.
  • FPR evaluators: NoiseFalsePositive, DiscoveryCountCap, MinValidatedDiscovery.
  • On-demand spec: full_profile_discrimination.spec.ts — live-retrieval with forbidden noise alert IDs and discovery count cap.
  • Skill hardening (attack_discovery_generator_skill.ts): routing keywords, ALERT_RETRIEVAL_BOUNDARIES (minimal get_default_esql_query → execute_esql path), forbidden generate_esql / alert-retrieval-builder misroutes, status-only get_status path.

Local A/B findings — skill hardening vs main baseline (2026-07-15)

Controlled comparison on the same harness (evals_attack_discovery_agent_builder, EIS Claude 4.6 Sonnet, --grep "golden |non-golden", rep=1):

Baseline (elastic#277625 @ 208d3c253b0a, main skill) This PR (d288ae72335f, hardened skill)
Playwright 5/5 (11.7m) 5/5 (11.9m)
CI on baseline branch #469078 PASS —

Evaluator Δ (golden-path means)

Dataset Metric Baseline Hardened Δ
live-retrieval Criteria 0 1 +1
Rubric 0.50 1 +0.50
ForbiddenTools 0 1 +1
Tool calls 11.5 4 −7.5
Trajectory 0.20 0.83 +0.63
WorkflowEvidence — 1 added
AdToolResult 0 0.50 +0.50 (still partial)
multiple-alert-sets Criteria 0 1 +1
Rubric 0 1 +1
AdToolResult 0 1 +1
WorkflowEvidence 0 1 +1
Tool calls 9 2 −7
Trajectory 0.14 1 +0.86
provided-alerts Tool calls 4 2 −2
Trajectory 0.33 1 +0.67
missing-alert-retrieval Basic 1 0 −1 (investigate)
Tool calls 9 3 −6
Latency 72s 10s −62s

Takeaway: Skill hardening fixes the two baseline weak spots (elastic#277625 §8) — live-retrieval misroutes/extra tooling and multiple-alert-sets quality gates — without breaking Playwright pass rate. missing-alert-retrieval Basic dropped 1→0 on one rep; worth a focused rerun before claiming regression.

How to run

Golden-path (weekly gate):

node scripts/evals run --suite attack-discovery-agent-builder --grep "golden |non-golden"

Full profile (on-demand):

node scripts/evals run --suite attack-discovery-agent-builder --grep "full profile"

Test plan

Extend the AD2 scenario registry with portable-seeder full profile seeding
(seven signal chains plus deterministic background and Defender noise), add
FPR evaluators (NoiseFalsePositive, DiscoveryCountCap, MinValidatedDiscovery),
and register an on-demand live-retrieval spec — not wired to weekly CI.
Move routing prose improvements out of the baseline eval PR so this
follow-up can be scored against elastic#277625 main-baseline skill behavior.
Finding #7: Add trigger phrases to SKILL_DESCRIPTION so TF-IDF routing
and weaker models match paraphrases like 'run across these alert sets'.

Finding #5: Move Field Syntax before Mode C and Missed Detection Closure
in SKILL_CONTENT so the core generate/status instructions appear first,
giving smaller models more usable context window before the appendix-like
post-report and gate sections.

Finding #8: Prefer insights from the run tool's attack_discoveries field
(parseInsightsFromToolResult) over regex-parsing the last ```json block
from the message. Falls back to message parsing with backwards search
for the insights key, fixing the case where a model emits a proposed
ES|QL rule after the insights JSON block.
patrykkopycinski pushed a commit that referenced this pull request Aug 5, 2026
## Summary

Set `connect.timeout = 60s` on the undici `Agent` used by
`KbnClientRequester` (https path only).

## Why

elastic#268531 migrated `KbnClient` from axios to native fetch but did not
override undici's 10s `connect.timeout` default. Axios had no equivalent
cutoff, so FTR callers talking to a busy local Kibana started failing
once that PR landed.

The `kibana-streams-performance` weekly pipeline went red in builds #9,
#11, #12, and #13 with:

```
ConnectTimeoutError: Connect Timeout Error (attempted address: localhost:5620, timeout: 10000ms)
```

The `10000ms` is undici's default. Bisect: build #8 last green
(2026-05-11) → #9 first red (2026-05-18), with elastic#268531 in the window.

## What changed


`src/platform/packages/shared/kbn-kbn-client/src/kbn_client/kbn_client_requester.ts`:
one constant, one option on the https `Agent`. http branch unchanged.

## Related

Regression introduced in elastic#268531. Companion streams perf PR: elastic#270636.

## Validation

https://buildkite.com/elastic/kibana-streams-performance/builds/14
patrykkopycinski added a commit to elastic/kibana that referenced this pull request Aug 11, 2026
## Summary

This PR introduces a dedicated, isolated evaluation suite for the
**Attack Discovery 2.0 Agent Builder** integration. It is intentionally
separate from the legacy direct-generation cohort so that we can measure
workflow, routing, tool efficiency, and output quality independently.

**The point of this PR is to show what focused evals buy you:** we did
not just add tests — running them surfaced routing noise, wasted tool
calls, and trace routing confusion, and those findings drove the **skill
and harness** changes included here.

**Baseline scope:** This PR ships the **eval suite + weekly CI wiring
only**. `attack-discovery-generator` skill prose is intentionally left
at **upstream/main** so follow-up PRs can measure improvement against
this baseline.

Committed evals use the **default Agent Builder router** — no
`configuration_overrides`. Dataset `expectedSkills` is scoring-only
(Skill Invoked evaluator).

---

## Eval profiles and CI cadence

This suite ships **two automated cohorts** in one package; they answer
different questions and run on different cadences.

| Profile | Spec | Seed | CI cadence | Primary question |
| --- | --- | --- | --- | --- |
| **Golden-path** | `attack_discovery_agent_builder.spec.ts` |
`fixtures.ts` (2 marker alerts) | **Weekly** — `llm_evals.yml` sets
`EVAL_GREP: 'golden |non-golden'` | Default-agent **routing**,
**tool/workflow plumbing**, efficiency gates |
| **Clean profile** | `clean_profile_provided_alerts.spec.ts` |
`scenario_registry/` (4 chains, 16 alerts + raw events) | **On-demand**
— run with `--grep "clean profile"` or full suite locally/Buildkite
on-demand | **Discovery quality** on realistic multi-stage chains
(portable-seeder `clean` parity) |
| **Full profile** | — | `ad-2.0-portable-seeder.py --profile full` |
**Not automated** in this PR | Signal vs **noise** (~150+ distractor
alerts); needs FPR/discrimination evaluators first |

**Do not merge golden-path and clean profile for weekly CI** —
golden-path includes live-retrieval, missing-retrieval, and status-only
cases that depend on the minimal marker fixture; clean profile is
provided-alerts-only on richer chain data.



## Evidence collected

### 1. Baseline (pre-prompt-tightening, first AD2 golden runs)

| Path | Tool calls | Input tokens | Notes |
|------|-----------|--------------|-------|
| **provided-alerts** | 5 | ~266k | Skill invoked, AD tool ran |
| **live-retrieval** | 13 | ~721k | Extra retrieval/corroboration tools
|

**Problem:** live-retrieval was doing 2.6× the tool calls and 2.7× the
tokens for the same 2-alert fixture. Quality was fine, but efficiency
was not.

### 2. After prompt tightening + new evaluators (clean stack restart)

| Path | Tool calls | Input tokens | AdToolResult | Basic | Criteria |
Rubric | Skill Invoked | Workflow | Trajectory |

|------|-----------|--------------|--------------|-------|----------|--------|---------------|----------|------------|
| **provided-alerts** | **4** | ~166k | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| **live-retrieval** | **10** | ~387k | 1 | 1 | 1 | 1 | 1 | 1 | ~0.30 |

**Improvement vs baseline:**

| Metric | provided-alerts | live-retrieval |
|--------|-----------------|----------------|
| Tool calls | 5 → **4** (−20%) | 13 → **10** (−23%) |
| Input tokens | ~266k → **~166k** (−38%) | ~721k → **~387k** (−46%) |
| Quality evaluators | already 1.0 | already 1.0 |

Trajectory on live-retrieval is intentionally low (~0.30) because the
agent legitimately runs `get_default_esql_query` → `execute_esql` →
`run`, but still adds corroboration detours. The new `ForbiddenTools`
and `CostPerAlert` evaluators make those detours visible instead of
hiding behind a quality score of 1.

### 3. Live-retrieval variance study (3 consecutive repetitions)

Even after tightening, live-retrieval is **non-deterministic**:

| Rep | Tool calls | Input tokens | Latency | ForbiddenTools |
CostPerAlert | Quality |

|-----|-----------|--------------|---------|----------------|--------------|---------|
| 1 | 9 | 413k | 37s | 0 | 1 | 1 |
| 2 | 14 | 657k | 135s | 1 | 0 | 1 |
| 3 | 11 | 586k | 38s | 0 | 0 | 1 |
| **mean** | **11.3** | **552k** | **70s** | **0.33** | **0.67** |
**1.0** |

- **ForbiddenTools** still leaks occasionally
(`platform.core.get_document_by_id`, `platform.core.generate_esql`).
- **CostPerAlert** fails 2/3 runs (threshold: 5 calls/alert × 2 alerts =
10; rep 2 hit 14).
- **Quality stays at 1.0** across all reps — efficiency is the remaining
gap, not output correctness.

### 4. What the evals revealed (and what we changed because of it)

| Finding | Evidence | Change in this PR |
|---------|----------|-------------------|
| Default router sometimes picks wrong skill | Live-retrieval misroutes
(baseline skill prose) | **Follow-up PR** — skill hardening moved out so
this PR stays baseline |
| Speculative tool calls inflate cost | 9–14 calls for 2 alerts;
`generate_esql` / `get_document_by_id` detours | **Follow-up PR** —
skill prompt tightening (baseline measures current behavior) |
| Trace evaluators returned `n/a` | Agent Builder spans on local Scout
ES, not golden cluster | Suite overrides `traceEsClient` to local Scout
ES |
| EIS connector setup flaky | Dot-prefixed `inferenceId` rejected;
manual cleanup between runs | Documented as operator setup issue —
**not** fixed in this PR (see out of scope) |
| Quality-only scoring hides waste | Rubric = 1 while trajectory = 0.30
and CostPerAlert = 0 | Added `ForbiddenTools`, `CostPerAlert`, strict
`Trajectory` evaluators |
| Only happy paths tested | Two golden cases insufficient for regression
signal | Added non-golden golden-path cases; clean-profile quality
cohort kept **on-demand** (not weekly) |

---

## What changed

### New evals package


`x-pack/solutions/security/packages/kbn-evals-suite-attack-discovery-agent-builder`

- Deterministic fixtures with **relative timestamps** (static
`@timestamp` caused live-retrieval to fail once alerts fell outside the
"last six hours" window).
- Custom evaluators: `WorkflowEvidence`, `StrictTrajectory`,
`ResponseSkillInvocation`, `ForbiddenTools`, `CostPerAlert`,
`AttackDiscoveryBasic`, `Criteria`, `Rubric`.
- `chat_client.ts` uses natural routing only — no harness
`configuration_overrides`; `expectedSkills` feeds Skill Invoked scoring.
- Trajectory precision excludes `load_skill` (framework routing cost
under natural routing).

### Dataset coverage

| Path | Type | What it validates |
|------|------|-------------------|
| `provided-alerts` | Golden | Alerts attached; agent skips retrieval |
| `live-retrieval` | Golden | Agent retrieves via default ES\|QL query |
| `multiple-alert-sets` | Non-golden | Multiple alert attachments in one
request |
| `missing-alert-retrieval` | Non-golden | Live retrieval returns zero
alerts; agent stops cleanly |
| `status-only` | Non-golden | User asks for `execution_uuid` status; no
new generation |
| `clean-profile` (scenario registry) | On-demand | Four portable-seeder
chains; **quality** evals (not weekly gate) |

### Infrastructure

- Scout config set: `evals_attack_discovery_agent_builder` (AD 2.0
feature flags).
- Suite registered in `evals.suites.json`, `package.json`,
`tsconfig.base.json`.
- Weekly pipeline: `llm_evals.yml` step for
`attack-discovery-agent-builder` with `EVAL_GREP` (golden-path only).
Clean profile remains in-repo for on-demand runs.

### Product improvements driven by evals

- **None in this PR** — skill-side routing prose changes moved to the
follow-up PR so this merge establishes a reproducible **main-baseline**
scorecard.

### Intentionally out of scope

This PR stays focused on the **AD2 eval suite**, **skill-side routing
prose**, and **weekly/on-demand CI registration**. The following were
explored during development but **reverted** to avoid unrelated platform
diffs:

- **`platform.core.generate_esql` tool description** — prefer-a-skill
boundary clause is a separate Agent Builder platform change, not part of
this evals PR.
- **`evals_tracing` EIS `inferenceId` normalization** — connector-cache
dot-prefix is an operator/local-setup concern; no Scout config change
ships here.
- **`@kbn/evals` `getToolCallSteps` `params` export** — the suite reads
`load_skill` arguments via local `getToolCallStepsWithParams()` instead
of extending the shared helper.

---

## How to run

```bash
node scripts/evals run --suite attack-discovery-agent-builder --connector eis-anthropic-claude-5-sonnet
```

Focused development:

```bash
node scripts/evals run --suite attack-discovery-agent-builder --grep "golden provided-alerts" --repetitions 1
```

---

## Follow-up (not blocking this PR)

- Further reduce live-retrieval variance (currently 9–14 tool calls;
CostPerAlert threshold is 5/alert).
- Scout ES stack stability under consecutive heavy runs (infrastructure,
not eval logic).
- Rubric judge flake on golden encoded-powershell export (`judge_failed`
/ orphaned `tool_result` — quality evaluators still pass).

## Related

- Skill-dev plugin lessons captured in
[elastic/agent-builder-skill-dev-cursor-plugin#149](elastic/agent-builder-skill-dev-cursor-plugin#149)
(`agent-builder-eval-playbook` skill).
- Does **not** merge with the legacy direct-generation cohort
(`kbn-evals-suite-attack-discovery`).

### 5. Weekly-matrix multi-model comparison (baseline vs after)

Local batch runs on **2026-07-13** (M4, `feat/search-skill-boundary`
worktree). LLM-gateway models skipped (`AD2_SKIP_GATEWAY_MODELS=true`).

| Cohort | TEST_RUN_ID | Playwright result |
|--------|-------------|-------------------|
| baseline | `ad2-baseline-20260713-112322` | **6/6 models — 5/5 tests
each** (`upstream/main` skill + `generate_esql`) |
| after (initial) | `ad2-after-20260713-144634` | **3/6 complete**
(Sonnet, Opus, Flash); 3 infra-failed |
| after (retry) | `ad2-after-retry-20260713-153105` | **1/3 recovered**
(GPT-OSS 5/5); Gemini 3.1 Pro + GPT-5.4 still infra-failed |

#### Infrastructure fix applied before retry

Dead-stack recovery was steering Playwright at `:5620` because
`.scout/servers/local.json` was clobbered by competing Scout boots while
the batch worker stayed on `:15001`/`:15000`. Batch runner now:

1. **`sync_scout_local_json()`** — patches (or recreates) `local.json`
hosts before every suite dispatch and after every boot/reboot.
2. **`ensure_stack_alive()`** — uses full two-pass `boot_stack()` on
dead-stack detection (not single-pass `boot_stack_once`), so EIS
connectors are re-baked after ES wipe.

#### Playwright pass matrix (merged after cohort)

| Model | Baseline | After (merged) | Notes |
|-------|----------|----------------|-------|
| Claude 4.6 Sonnet | 5/5 | **5/5** | after initial run |
| Claude 4.6 Opus | 5/5 | **5/5** | after initial run |
| Gemini 3.0 Flash | 5/5 | **5/5** | after initial run |
| Gemini 3.1 Pro | 5/5 | **2/5 fail** (retry also 2/5) | Stack died on
`:15001` during converse; later tests hit missing `local.json` |
| GPT-5.4 | 5/5 | **0/5** | Connector race at `playwright.config.ts`
load (`Evaluation connector id … was not found`) — both initial + retry
|
| GPT-OSS-120b | 5/5 | **5/5** | Retry succeeded on `:15001`/`:15000`
after port-sync fix |

#### Evaluator Δ table (Overall means, log-parsed)

Scores extracted from Playwright summary tables in batch logs (ephemeral
Scout ES on `:15000` was torn down before ES-indexed export;
`generate-ad2-comparison-report.mjs` could not query
`.evaluation-scores*` post-teardown). **Δ = after − baseline.**

| Model | PW | AdToolResult | WorkflowEvidence | trajectory |
ForbiddenTools | AttackDiscoveryBasic | Criteria | Rubric | Skill
Invoked | Tool Calls | Latency |

|-------|-----|-------------:|-----------------:|-----------:|---------------:|---------------------:|---------:|-------:|--------------:|-----------:|--------:|
| Claude 4.6 Sonnet | 5/5→5/5 | 0.50→0.50 (0) | 0.83→0.67 (−0.16) | 1→1
(0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 0.50→0.52 (+0.02) |
6.7→6.3 | 34.2s→31.4s |
| Claude 4.6 Opus | 5/5→5/5 | 0.83→0.67 (−0.16) | 0.50→0.83 (+0.33) |
0.80→0.75 (−0.05) | 0.33→1 (+0.67) | 0.50→0.67 (+0.17) | 1→1 (0) |
0.33→0.83 (+0.50) | 0.24→0.37 (+0.13) | 18.3→12.0 | 60.5s→42.2s |
| Gemini 3.0 Flash | 5/5→5/5 | 0.50→0.50 (0) | 0.50→0.50 (0) | 0.67→0.33
(−0.34) | 0.93→0.60 (−0.33) | 1→0.67 (−0.33) | 1→1 (0) | 1→0.67 (−0.33)
| 0.16→0.23 (+0.07) | 15.2→21.2 | 13.2s→19.6s |
| Gemini 3.1 Pro | 5/5→**2/5 fail** | — | — | — | — | — | — | — | — | —
| — |
| GPT-5.4 | 5/5→**0/5** | — | — | — | — | — | — | — | — | — | — |
| GPT-OSS-120b | 5/5→5/5 | 0.50→0.00 (−0.50) | 0.17→0.17 (0) | 1→— |
0.33→0.40 (+0.07) | 1→0.83 (−0.17) | 1→0.83 (−0.17) | 0.67→0.83 (+0.16)
| 0.43→0.00 (−0.43) | 5.2→12.2 | 52.9s→30.4s |

Baseline GPT-OSS row from `ad2-baseline-20260713-112322` run log; after
GPT-OSS row from retry run `ad2-after-retry-20260713-153105`.

#### Regression watch (completed models only)

On the **4 models with full 5/5 after Playwright completion** (Sonnet,
Opus, Flash, GPT-OSS):

- **No Criteria / AttackDiscoveryBasic regressions** — Criteria stays at
1.0 where scored; Basic dips on Flash (−0.33) and GPT-OSS (−0.17) but
remains ≥0.67.
- **Efficiency mixed:** Opus improves (fewer tool calls, lower latency);
Flash regresses on trajectory (−0.34) and tool calls (+6).
- **GPT-OSS quality metrics look lower** (AdToolResult 0, Skill Invoked
0) but this matches baseline's already-low AdTool/Skill scores — not
introduced by after-cohort hardening.

**Cannot conclude** for Gemini 3.1 Pro or GPT-5.4 — failures are
stack/connector infrastructure, not scored quality regressions.

#### Harness notes

- Stop competing Scout stacks before batch: `node scripts/evals stop`
- `TEST_KIBANA_PORT`/`TEST_ES_PORT` patch in `base.config.ts` required
when branch lacks `fix/weekly-evals-matrix` merge
- Batch runner port-sync + full `boot_stack()` recovery landed in
`run-security-evals-batch.sh` (skill-dev plugin, local)
- Remaining gap: per-model connector-registration wait before
`playwright.config.ts` loads (GPT-5.4 race)

### 6. Natural-routing convergence + OTLP trace evidence
(`ad2-full-natural-20260713-191333`)

**Harness change:** removed `configuration_overrides` /
`AD2_EVAL_FORCE_SKILL_ROUTING` — default Agent Builder router only;
`expectedSkills` is scoring-only.

**Run:** `ad2-full-natural-20260713-191333` on
`feat/search-skill-boundary` worktree (Claude 4.6 Sonnet, local Scout
`:15001`/`:15000`). **9/9 Playwright passed** (25.3m).

#### Representative trace cases (OTLP on local Scout ES — attach before
stack restart)

| Case | trace.id | Spans | Skill Invoked | ForbiddenTools | trajectory
| Tool Calls | Quality |

|------|----------|-------|---------------|----------------|------------|------------|---------|
| **live-retrieval** | `bf1f541ffa9a8e91af088dd526a74ae6` | 13 | **1.0**
| **1.0** | 0.75 | 4 | Criteria/Rubric/AdToolResult **1.0** |
| **clean-profile (encoded-powershell)** |
`f49eca1730170cf49fbf48b6a89fca0e` | 9 | **1.0** | **1.0** | 0.50 | 2 |
Criteria/Rubric/AdToolResult **1.0** |

**Live-retrieval OTLP tool sequence:** `load_skill` →
`get_default_esql_query` (custom) → `platform.core.execute_esql` →
`security.attack-discovery.run` (custom). No `generate_esql`, no
`attack-discovery-alert-retrieval-builder`.

**Clean-profile OTLP tool sequence:** `load_skill` →
`security.attack-discovery.run` with provided alert IDs.

**Skill-side fix (routing):** `attack_discovery_generator_skill.ts` —
routing keywords, `ALERT_RETRIEVAL_BOUNDARIES` (minimal
`get_default_esql_query` → `execute_esql` path), `STATUS_GUIDE` for
status-only queries.

**Local evidence artifacts** (operator worktree, gitignored):
`.ad2-pr-evidence/ad2-full-natural-20260713-191333/` — `README.md`,
`scores/score-summary.json`, `traces/*-otlp-projected.json`, raw OTLP
JSON.

### Golden-cluster OTLP trace IDs (exported 2026-07-13)

| Case | Golden `trace.id` | Run ID | Skill Invoked | ForbiddenTools |
Waterfall |

|------|-------------------|--------|---------------|----------------|-----------|
| **live-retrieval** (happy path) | `7c29a6966e8766ffe343c9e3f8516aeb` |
`ad2-golden-export-live-20260713-204938` | 1.0 | 1.0 | [APM
trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=7c29a6966e8766ffe343c9e3f8516aeb)
|
| **clean-profile encoded-powershell** |
`594f040b9f47df82f955107e6aaf5019` | `ad2-golden-export-20260713-203039`
| 1.0 | 1.0 | [APM
trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=594f040b9f47df82f955107e6aaf5019)
|

Local Scout OTLP snapshots (pre-teardown):
`bf1f541ffa9a8e91af088dd526a74ae6`, `f49eca1730170cf49fbf48b6a89fca0e` —
see `.ad2-pr-evidence/ad2-full-natural-20260713-191333/`.

### 7. Post-merge natural routing (2026-07-13)

Latest commit on this branch drops harness forcing, adds clean-profile
scenario-registry evals, and tunes trajectory precision to exclude
`load_skill` under natural routing.
### 8. Skill-baseline golden-path scorecard (2026-07-15)

**CI:** [Buildkite
#469078](https://buildkite.com/elastic/kibana-pull-request/builds/469078)
**PASS** on `208d3c253b0a` (skill prose reverted to upstream/main).

**Local run:** `feat/attack-discovery-agent-builder-evals` @
`208d3c253b0a`, EIS Claude 4.6 Sonnet, `--grep "golden |non-golden"`,
rep=1, Scout `evals_attack_discovery_agent_builder`. **5/5 Playwright
passed** (11.7m).

| Dataset | AdToolResult | Basic | Criteria | ForbiddenTools | Rubric |
Skill Invoked | Tool Calls | WorkflowEvidence | Trajectory |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| provided-alerts | 1 | 1 | 1 | 1 | 1 | 1 | 4 | 1 | 0.33 |
| live-retrieval (n=2) | 0 | 0.50 | **0** | **0** | 0.50 | 1 | **11.5**
| — | **0.20** |
| multiple-alert-sets | 0 | 1 | **0** | 1 | **0** | 1 | **9** | **0** |
0.14 |
| missing-alert-retrieval | 0 | 1 | — | 0 | 1 | 1 | 9 | — | 0.29 |
| status-only | 0 | 0 | — | 1 | 1 | 1 | 2 | — | 1 |

**Interpretation:** With **main skill prose**, routing works (Skill
Invoked = 1 everywhere) but **live-retrieval** and
**multiple-alert-sets** show efficiency/quality gaps — extra tool calls,
low Criteria/Rubric, missing WorkflowEvidence. These are the comparison
targets for skill hardening in follow-up
[patrykkopycinski#13](patrykkopycinski#13)
(not in this PR).

**Follow-up A/B (skill overlay from #13, same harness):** live-retrieval
Criteria 0→1, Rubric 0.50→1, ForbiddenTools 0→1, tool calls 11.5→4,
trajectory 0.20→0.83; multiple-alert-sets Criteria 0→1, Rubric 0→1,
WorkflowEvidence 0→1, tool calls 9→2. See #13 for full Δ table.

---------

Signed-off-by: Patryk Kopycinski <patryk.kopycinski@elastic.co>
Co-authored-by: kibanamachine <42973632+kibanamachine@users.noreply.github.com>
qn895 pushed a commit to qn895/kibana that referenced this pull request Aug 11, 2026
…277625)

## Summary

This PR introduces a dedicated, isolated evaluation suite for the
**Attack Discovery 2.0 Agent Builder** integration. It is intentionally
separate from the legacy direct-generation cohort so that we can measure
workflow, routing, tool efficiency, and output quality independently.

**The point of this PR is to show what focused evals buy you:** we did
not just add tests — running them surfaced routing noise, wasted tool
calls, and trace routing confusion, and those findings drove the **skill
and harness** changes included here.

**Baseline scope:** This PR ships the **eval suite + weekly CI wiring
only**. `attack-discovery-generator` skill prose is intentionally left
at **upstream/main** so follow-up PRs can measure improvement against
this baseline.

Committed evals use the **default Agent Builder router** — no
`configuration_overrides`. Dataset `expectedSkills` is scoring-only
(Skill Invoked evaluator).

---

## Eval profiles and CI cadence

This suite ships **two automated cohorts** in one package; they answer
different questions and run on different cadences.

| Profile | Spec | Seed | CI cadence | Primary question |
| --- | --- | --- | --- | --- |
| **Golden-path** | `attack_discovery_agent_builder.spec.ts` |
`fixtures.ts` (2 marker alerts) | **Weekly** — `llm_evals.yml` sets
`EVAL_GREP: 'golden |non-golden'` | Default-agent **routing**,
**tool/workflow plumbing**, efficiency gates |
| **Clean profile** | `clean_profile_provided_alerts.spec.ts` |
`scenario_registry/` (4 chains, 16 alerts + raw events) | **On-demand**
— run with `--grep "clean profile"` or full suite locally/Buildkite
on-demand | **Discovery quality** on realistic multi-stage chains
(portable-seeder `clean` parity) |
| **Full profile** | — | `ad-2.0-portable-seeder.py --profile full` |
**Not automated** in this PR | Signal vs **noise** (~150+ distractor
alerts); needs FPR/discrimination evaluators first |

**Do not merge golden-path and clean profile for weekly CI** —
golden-path includes live-retrieval, missing-retrieval, and status-only
cases that depend on the minimal marker fixture; clean profile is
provided-alerts-only on richer chain data.



## Evidence collected

### 1. Baseline (pre-prompt-tightening, first AD2 golden runs)

| Path | Tool calls | Input tokens | Notes |
|------|-----------|--------------|-------|
| **provided-alerts** | 5 | ~266k | Skill invoked, AD tool ran |
| **live-retrieval** | 13 | ~721k | Extra retrieval/corroboration tools
|

**Problem:** live-retrieval was doing 2.6× the tool calls and 2.7× the
tokens for the same 2-alert fixture. Quality was fine, but efficiency
was not.

### 2. After prompt tightening + new evaluators (clean stack restart)

| Path | Tool calls | Input tokens | AdToolResult | Basic | Criteria |
Rubric | Skill Invoked | Workflow | Trajectory |

|------|-----------|--------------|--------------|-------|----------|--------|---------------|----------|------------|
| **provided-alerts** | **4** | ~166k | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| **live-retrieval** | **10** | ~387k | 1 | 1 | 1 | 1 | 1 | 1 | ~0.30 |

**Improvement vs baseline:**

| Metric | provided-alerts | live-retrieval |
|--------|-----------------|----------------|
| Tool calls | 5 → **4** (−20%) | 13 → **10** (−23%) |
| Input tokens | ~266k → **~166k** (−38%) | ~721k → **~387k** (−46%) |
| Quality evaluators | already 1.0 | already 1.0 |

Trajectory on live-retrieval is intentionally low (~0.30) because the
agent legitimately runs `get_default_esql_query` → `execute_esql` →
`run`, but still adds corroboration detours. The new `ForbiddenTools`
and `CostPerAlert` evaluators make those detours visible instead of
hiding behind a quality score of 1.

### 3. Live-retrieval variance study (3 consecutive repetitions)

Even after tightening, live-retrieval is **non-deterministic**:

| Rep | Tool calls | Input tokens | Latency | ForbiddenTools |
CostPerAlert | Quality |

|-----|-----------|--------------|---------|----------------|--------------|---------|
| 1 | 9 | 413k | 37s | 0 | 1 | 1 |
| 2 | 14 | 657k | 135s | 1 | 0 | 1 |
| 3 | 11 | 586k | 38s | 0 | 0 | 1 |
| **mean** | **11.3** | **552k** | **70s** | **0.33** | **0.67** |
**1.0** |

- **ForbiddenTools** still leaks occasionally
(`platform.core.get_document_by_id`, `platform.core.generate_esql`).
- **CostPerAlert** fails 2/3 runs (threshold: 5 calls/alert × 2 alerts =
10; rep 2 hit 14).
- **Quality stays at 1.0** across all reps — efficiency is the remaining
gap, not output correctness.

### 4. What the evals revealed (and what we changed because of it)

| Finding | Evidence | Change in this PR |
|---------|----------|-------------------|
| Default router sometimes picks wrong skill | Live-retrieval misroutes
(baseline skill prose) | **Follow-up PR** — skill hardening moved out so
this PR stays baseline |
| Speculative tool calls inflate cost | 9–14 calls for 2 alerts;
`generate_esql` / `get_document_by_id` detours | **Follow-up PR** —
skill prompt tightening (baseline measures current behavior) |
| Trace evaluators returned `n/a` | Agent Builder spans on local Scout
ES, not golden cluster | Suite overrides `traceEsClient` to local Scout
ES |
| EIS connector setup flaky | Dot-prefixed `inferenceId` rejected;
manual cleanup between runs | Documented as operator setup issue —
**not** fixed in this PR (see out of scope) |
| Quality-only scoring hides waste | Rubric = 1 while trajectory = 0.30
and CostPerAlert = 0 | Added `ForbiddenTools`, `CostPerAlert`, strict
`Trajectory` evaluators |
| Only happy paths tested | Two golden cases insufficient for regression
signal | Added non-golden golden-path cases; clean-profile quality
cohort kept **on-demand** (not weekly) |

---

## What changed

### New evals package


`x-pack/solutions/security/packages/kbn-evals-suite-attack-discovery-agent-builder`

- Deterministic fixtures with **relative timestamps** (static
`@timestamp` caused live-retrieval to fail once alerts fell outside the
"last six hours" window).
- Custom evaluators: `WorkflowEvidence`, `StrictTrajectory`,
`ResponseSkillInvocation`, `ForbiddenTools`, `CostPerAlert`,
`AttackDiscoveryBasic`, `Criteria`, `Rubric`.
- `chat_client.ts` uses natural routing only — no harness
`configuration_overrides`; `expectedSkills` feeds Skill Invoked scoring.
- Trajectory precision excludes `load_skill` (framework routing cost
under natural routing).

### Dataset coverage

| Path | Type | What it validates |
|------|------|-------------------|
| `provided-alerts` | Golden | Alerts attached; agent skips retrieval |
| `live-retrieval` | Golden | Agent retrieves via default ES\|QL query |
| `multiple-alert-sets` | Non-golden | Multiple alert attachments in one
request |
| `missing-alert-retrieval` | Non-golden | Live retrieval returns zero
alerts; agent stops cleanly |
| `status-only` | Non-golden | User asks for `execution_uuid` status; no
new generation |
| `clean-profile` (scenario registry) | On-demand | Four portable-seeder
chains; **quality** evals (not weekly gate) |

### Infrastructure

- Scout config set: `evals_attack_discovery_agent_builder` (AD 2.0
feature flags).
- Suite registered in `evals.suites.json`, `package.json`,
`tsconfig.base.json`.
- Weekly pipeline: `llm_evals.yml` step for
`attack-discovery-agent-builder` with `EVAL_GREP` (golden-path only).
Clean profile remains in-repo for on-demand runs.

### Product improvements driven by evals

- **None in this PR** — skill-side routing prose changes moved to the
follow-up PR so this merge establishes a reproducible **main-baseline**
scorecard.

### Intentionally out of scope

This PR stays focused on the **AD2 eval suite**, **skill-side routing
prose**, and **weekly/on-demand CI registration**. The following were
explored during development but **reverted** to avoid unrelated platform
diffs:

- **`platform.core.generate_esql` tool description** — prefer-a-skill
boundary clause is a separate Agent Builder platform change, not part of
this evals PR.
- **`evals_tracing` EIS `inferenceId` normalization** — connector-cache
dot-prefix is an operator/local-setup concern; no Scout config change
ships here.
- **`@kbn/evals` `getToolCallSteps` `params` export** — the suite reads
`load_skill` arguments via local `getToolCallStepsWithParams()` instead
of extending the shared helper.

---

## How to run

```bash
node scripts/evals run --suite attack-discovery-agent-builder --connector eis-anthropic-claude-5-sonnet
```

Focused development:

```bash
node scripts/evals run --suite attack-discovery-agent-builder --grep "golden provided-alerts" --repetitions 1
```

---

## Follow-up (not blocking this PR)

- Further reduce live-retrieval variance (currently 9–14 tool calls;
CostPerAlert threshold is 5/alert).
- Scout ES stack stability under consecutive heavy runs (infrastructure,
not eval logic).
- Rubric judge flake on golden encoded-powershell export (`judge_failed`
/ orphaned `tool_result` — quality evaluators still pass).

## Related

- Skill-dev plugin lessons captured in
[elastic/agent-builder-skill-dev-cursor-plugin#149](elastic/agent-builder-skill-dev-cursor-plugin#149)
(`agent-builder-eval-playbook` skill).
- Does **not** merge with the legacy direct-generation cohort
(`kbn-evals-suite-attack-discovery`).

### 5. Weekly-matrix multi-model comparison (baseline vs after)

Local batch runs on **2026-07-13** (M4, `feat/search-skill-boundary`
worktree). LLM-gateway models skipped (`AD2_SKIP_GATEWAY_MODELS=true`).

| Cohort | TEST_RUN_ID | Playwright result |
|--------|-------------|-------------------|
| baseline | `ad2-baseline-20260713-112322` | **6/6 models — 5/5 tests
each** (`upstream/main` skill + `generate_esql`) |
| after (initial) | `ad2-after-20260713-144634` | **3/6 complete**
(Sonnet, Opus, Flash); 3 infra-failed |
| after (retry) | `ad2-after-retry-20260713-153105` | **1/3 recovered**
(GPT-OSS 5/5); Gemini 3.1 Pro + GPT-5.4 still infra-failed |

#### Infrastructure fix applied before retry

Dead-stack recovery was steering Playwright at `:5620` because
`.scout/servers/local.json` was clobbered by competing Scout boots while
the batch worker stayed on `:15001`/`:15000`. Batch runner now:

1. **`sync_scout_local_json()`** — patches (or recreates) `local.json`
hosts before every suite dispatch and after every boot/reboot.
2. **`ensure_stack_alive()`** — uses full two-pass `boot_stack()` on
dead-stack detection (not single-pass `boot_stack_once`), so EIS
connectors are re-baked after ES wipe.

#### Playwright pass matrix (merged after cohort)

| Model | Baseline | After (merged) | Notes |
|-------|----------|----------------|-------|
| Claude 4.6 Sonnet | 5/5 | **5/5** | after initial run |
| Claude 4.6 Opus | 5/5 | **5/5** | after initial run |
| Gemini 3.0 Flash | 5/5 | **5/5** | after initial run |
| Gemini 3.1 Pro | 5/5 | **2/5 fail** (retry also 2/5) | Stack died on
`:15001` during converse; later tests hit missing `local.json` |
| GPT-5.4 | 5/5 | **0/5** | Connector race at `playwright.config.ts`
load (`Evaluation connector id … was not found`) — both initial + retry
|
| GPT-OSS-120b | 5/5 | **5/5** | Retry succeeded on `:15001`/`:15000`
after port-sync fix |

#### Evaluator Δ table (Overall means, log-parsed)

Scores extracted from Playwright summary tables in batch logs (ephemeral
Scout ES on `:15000` was torn down before ES-indexed export;
`generate-ad2-comparison-report.mjs` could not query
`.evaluation-scores*` post-teardown). **Δ = after − baseline.**

| Model | PW | AdToolResult | WorkflowEvidence | trajectory |
ForbiddenTools | AttackDiscoveryBasic | Criteria | Rubric | Skill
Invoked | Tool Calls | Latency |

|-------|-----|-------------:|-----------------:|-----------:|---------------:|---------------------:|---------:|-------:|--------------:|-----------:|--------:|
| Claude 4.6 Sonnet | 5/5→5/5 | 0.50→0.50 (0) | 0.83→0.67 (−0.16) | 1→1
(0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 0.50→0.52 (+0.02) |
6.7→6.3 | 34.2s→31.4s |
| Claude 4.6 Opus | 5/5→5/5 | 0.83→0.67 (−0.16) | 0.50→0.83 (+0.33) |
0.80→0.75 (−0.05) | 0.33→1 (+0.67) | 0.50→0.67 (+0.17) | 1→1 (0) |
0.33→0.83 (+0.50) | 0.24→0.37 (+0.13) | 18.3→12.0 | 60.5s→42.2s |
| Gemini 3.0 Flash | 5/5→5/5 | 0.50→0.50 (0) | 0.50→0.50 (0) | 0.67→0.33
(−0.34) | 0.93→0.60 (−0.33) | 1→0.67 (−0.33) | 1→1 (0) | 1→0.67 (−0.33)
| 0.16→0.23 (+0.07) | 15.2→21.2 | 13.2s→19.6s |
| Gemini 3.1 Pro | 5/5→**2/5 fail** | — | — | — | — | — | — | — | — | —
| — |
| GPT-5.4 | 5/5→**0/5** | — | — | — | — | — | — | — | — | — | — |
| GPT-OSS-120b | 5/5→5/5 | 0.50→0.00 (−0.50) | 0.17→0.17 (0) | 1→— |
0.33→0.40 (+0.07) | 1→0.83 (−0.17) | 1→0.83 (−0.17) | 0.67→0.83 (+0.16)
| 0.43→0.00 (−0.43) | 5.2→12.2 | 52.9s→30.4s |

Baseline GPT-OSS row from `ad2-baseline-20260713-112322` run log; after
GPT-OSS row from retry run `ad2-after-retry-20260713-153105`.

#### Regression watch (completed models only)

On the **4 models with full 5/5 after Playwright completion** (Sonnet,
Opus, Flash, GPT-OSS):

- **No Criteria / AttackDiscoveryBasic regressions** — Criteria stays at
1.0 where scored; Basic dips on Flash (−0.33) and GPT-OSS (−0.17) but
remains ≥0.67.
- **Efficiency mixed:** Opus improves (fewer tool calls, lower latency);
Flash regresses on trajectory (−0.34) and tool calls (+6).
- **GPT-OSS quality metrics look lower** (AdToolResult 0, Skill Invoked
0) but this matches baseline's already-low AdTool/Skill scores — not
introduced by after-cohort hardening.

**Cannot conclude** for Gemini 3.1 Pro or GPT-5.4 — failures are
stack/connector infrastructure, not scored quality regressions.

#### Harness notes

- Stop competing Scout stacks before batch: `node scripts/evals stop`
- `TEST_KIBANA_PORT`/`TEST_ES_PORT` patch in `base.config.ts` required
when branch lacks `fix/weekly-evals-matrix` merge
- Batch runner port-sync + full `boot_stack()` recovery landed in
`run-security-evals-batch.sh` (skill-dev plugin, local)
- Remaining gap: per-model connector-registration wait before
`playwright.config.ts` loads (GPT-5.4 race)

### 6. Natural-routing convergence + OTLP trace evidence
(`ad2-full-natural-20260713-191333`)

**Harness change:** removed `configuration_overrides` /
`AD2_EVAL_FORCE_SKILL_ROUTING` — default Agent Builder router only;
`expectedSkills` is scoring-only.

**Run:** `ad2-full-natural-20260713-191333` on
`feat/search-skill-boundary` worktree (Claude 4.6 Sonnet, local Scout
`:15001`/`:15000`). **9/9 Playwright passed** (25.3m).

#### Representative trace cases (OTLP on local Scout ES — attach before
stack restart)

| Case | trace.id | Spans | Skill Invoked | ForbiddenTools | trajectory
| Tool Calls | Quality |

|------|----------|-------|---------------|----------------|------------|------------|---------|
| **live-retrieval** | `bf1f541ffa9a8e91af088dd526a74ae6` | 13 | **1.0**
| **1.0** | 0.75 | 4 | Criteria/Rubric/AdToolResult **1.0** |
| **clean-profile (encoded-powershell)** |
`f49eca1730170cf49fbf48b6a89fca0e` | 9 | **1.0** | **1.0** | 0.50 | 2 |
Criteria/Rubric/AdToolResult **1.0** |

**Live-retrieval OTLP tool sequence:** `load_skill` →
`get_default_esql_query` (custom) → `platform.core.execute_esql` →
`security.attack-discovery.run` (custom). No `generate_esql`, no
`attack-discovery-alert-retrieval-builder`.

**Clean-profile OTLP tool sequence:** `load_skill` →
`security.attack-discovery.run` with provided alert IDs.

**Skill-side fix (routing):** `attack_discovery_generator_skill.ts` —
routing keywords, `ALERT_RETRIEVAL_BOUNDARIES` (minimal
`get_default_esql_query` → `execute_esql` path), `STATUS_GUIDE` for
status-only queries.

**Local evidence artifacts** (operator worktree, gitignored):
`.ad2-pr-evidence/ad2-full-natural-20260713-191333/` — `README.md`,
`scores/score-summary.json`, `traces/*-otlp-projected.json`, raw OTLP
JSON.

### Golden-cluster OTLP trace IDs (exported 2026-07-13)

| Case | Golden `trace.id` | Run ID | Skill Invoked | ForbiddenTools |
Waterfall |

|------|-------------------|--------|---------------|----------------|-----------|
| **live-retrieval** (happy path) | `7c29a6966e8766ffe343c9e3f8516aeb` |
`ad2-golden-export-live-20260713-204938` | 1.0 | 1.0 | [APM
trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=7c29a6966e8766ffe343c9e3f8516aeb)
|
| **clean-profile encoded-powershell** |
`594f040b9f47df82f955107e6aaf5019` | `ad2-golden-export-20260713-203039`
| 1.0 | 1.0 | [APM
trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=594f040b9f47df82f955107e6aaf5019)
|

Local Scout OTLP snapshots (pre-teardown):
`bf1f541ffa9a8e91af088dd526a74ae6`, `f49eca1730170cf49fbf48b6a89fca0e` —
see `.ad2-pr-evidence/ad2-full-natural-20260713-191333/`.

### 7. Post-merge natural routing (2026-07-13)

Latest commit on this branch drops harness forcing, adds clean-profile
scenario-registry evals, and tunes trajectory precision to exclude
`load_skill` under natural routing.
### 8. Skill-baseline golden-path scorecard (2026-07-15)

**CI:** [Buildkite
#469078](https://buildkite.com/elastic/kibana-pull-request/builds/469078)
**PASS** on `208d3c253b0a` (skill prose reverted to upstream/main).

**Local run:** `feat/attack-discovery-agent-builder-evals` @
`208d3c253b0a`, EIS Claude 4.6 Sonnet, `--grep "golden |non-golden"`,
rep=1, Scout `evals_attack_discovery_agent_builder`. **5/5 Playwright
passed** (11.7m).

| Dataset | AdToolResult | Basic | Criteria | ForbiddenTools | Rubric |
Skill Invoked | Tool Calls | WorkflowEvidence | Trajectory |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| provided-alerts | 1 | 1 | 1 | 1 | 1 | 1 | 4 | 1 | 0.33 |
| live-retrieval (n=2) | 0 | 0.50 | **0** | **0** | 0.50 | 1 | **11.5**
| — | **0.20** |
| multiple-alert-sets | 0 | 1 | **0** | 1 | **0** | 1 | **9** | **0** |
0.14 |
| missing-alert-retrieval | 0 | 1 | — | 0 | 1 | 1 | 9 | — | 0.29 |
| status-only | 0 | 0 | — | 1 | 1 | 1 | 2 | — | 1 |

**Interpretation:** With **main skill prose**, routing works (Skill
Invoked = 1 everywhere) but **live-retrieval** and
**multiple-alert-sets** show efficiency/quality gaps — extra tool calls,
low Criteria/Rubric, missing WorkflowEvidence. These are the comparison
targets for skill hardening in follow-up
[patrykkopycinski#13](patrykkopycinski#13)
(not in this PR).

**Follow-up A/B (skill overlay from elastic#13, same harness):** live-retrieval
Criteria 0→1, Rubric 0.50→1, ForbiddenTools 0→1, tool calls 11.5→4,
trajectory 0.20→0.83; multiple-alert-sets Criteria 0→1, Rubric 0→1,
WorkflowEvidence 0→1, tool calls 9→2. See elastic#13 for full Δ table.

---------

Signed-off-by: Patryk Kopycinski <patryk.kopycinski@elastic.co>
Co-authored-by: kibanamachine <42973632+kibanamachine@users.noreply.github.com>
patrykkopycinski added a commit that referenced this pull request Aug 18, 2026
…277625)

## Summary

This PR introduces a dedicated, isolated evaluation suite for the
**Attack Discovery 2.0 Agent Builder** integration. It is intentionally
separate from the legacy direct-generation cohort so that we can measure
workflow, routing, tool efficiency, and output quality independently.

**The point of this PR is to show what focused evals buy you:** we did
not just add tests — running them surfaced routing noise, wasted tool
calls, and trace routing confusion, and those findings drove the **skill
and harness** changes included here.

**Baseline scope:** This PR ships the **eval suite + weekly CI wiring
only**. `attack-discovery-generator` skill prose is intentionally left
at **upstream/main** so follow-up PRs can measure improvement against
this baseline.

Committed evals use the **default Agent Builder router** — no
`configuration_overrides`. Dataset `expectedSkills` is scoring-only
(Skill Invoked evaluator).

---

## Eval profiles and CI cadence

This suite ships **two automated cohorts** in one package; they answer
different questions and run on different cadences.

| Profile | Spec | Seed | CI cadence | Primary question |
| --- | --- | --- | --- | --- |
| **Golden-path** | `attack_discovery_agent_builder.spec.ts` |
`fixtures.ts` (2 marker alerts) | **Weekly** — `llm_evals.yml` sets
`EVAL_GREP: 'golden |non-golden'` | Default-agent **routing**,
**tool/workflow plumbing**, efficiency gates |
| **Clean profile** | `clean_profile_provided_alerts.spec.ts` |
`scenario_registry/` (4 chains, 16 alerts + raw events) | **On-demand**
— run with `--grep "clean profile"` or full suite locally/Buildkite
on-demand | **Discovery quality** on realistic multi-stage chains
(portable-seeder `clean` parity) |
| **Full profile** | — | `ad-2.0-portable-seeder.py --profile full` |
**Not automated** in this PR | Signal vs **noise** (~150+ distractor
alerts); needs FPR/discrimination evaluators first |

**Do not merge golden-path and clean profile for weekly CI** —
golden-path includes live-retrieval, missing-retrieval, and status-only
cases that depend on the minimal marker fixture; clean profile is
provided-alerts-only on richer chain data.



## Evidence collected

### 1. Baseline (pre-prompt-tightening, first AD2 golden runs)

| Path | Tool calls | Input tokens | Notes |
|------|-----------|--------------|-------|
| **provided-alerts** | 5 | ~266k | Skill invoked, AD tool ran |
| **live-retrieval** | 13 | ~721k | Extra retrieval/corroboration tools
|

**Problem:** live-retrieval was doing 2.6× the tool calls and 2.7× the
tokens for the same 2-alert fixture. Quality was fine, but efficiency
was not.

### 2. After prompt tightening + new evaluators (clean stack restart)

| Path | Tool calls | Input tokens | AdToolResult | Basic | Criteria |
Rubric | Skill Invoked | Workflow | Trajectory |

|------|-----------|--------------|--------------|-------|----------|--------|---------------|----------|------------|
| **provided-alerts** | **4** | ~166k | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| **live-retrieval** | **10** | ~387k | 1 | 1 | 1 | 1 | 1 | 1 | ~0.30 |

**Improvement vs baseline:**

| Metric | provided-alerts | live-retrieval |
|--------|-----------------|----------------|
| Tool calls | 5 → **4** (−20%) | 13 → **10** (−23%) |
| Input tokens | ~266k → **~166k** (−38%) | ~721k → **~387k** (−46%) |
| Quality evaluators | already 1.0 | already 1.0 |

Trajectory on live-retrieval is intentionally low (~0.30) because the
agent legitimately runs `get_default_esql_query` → `execute_esql` →
`run`, but still adds corroboration detours. The new `ForbiddenTools`
and `CostPerAlert` evaluators make those detours visible instead of
hiding behind a quality score of 1.

### 3. Live-retrieval variance study (3 consecutive repetitions)

Even after tightening, live-retrieval is **non-deterministic**:

| Rep | Tool calls | Input tokens | Latency | ForbiddenTools |
CostPerAlert | Quality |

|-----|-----------|--------------|---------|----------------|--------------|---------|
| 1 | 9 | 413k | 37s | 0 | 1 | 1 |
| 2 | 14 | 657k | 135s | 1 | 0 | 1 |
| 3 | 11 | 586k | 38s | 0 | 0 | 1 |
| **mean** | **11.3** | **552k** | **70s** | **0.33** | **0.67** |
**1.0** |

- **ForbiddenTools** still leaks occasionally
(`platform.core.get_document_by_id`, `platform.core.generate_esql`).
- **CostPerAlert** fails 2/3 runs (threshold: 5 calls/alert × 2 alerts =
10; rep 2 hit 14).
- **Quality stays at 1.0** across all reps — efficiency is the remaining
gap, not output correctness.

### 4. What the evals revealed (and what we changed because of it)

| Finding | Evidence | Change in this PR |
|---------|----------|-------------------|
| Default router sometimes picks wrong skill | Live-retrieval misroutes
(baseline skill prose) | **Follow-up PR** — skill hardening moved out so
this PR stays baseline |
| Speculative tool calls inflate cost | 9–14 calls for 2 alerts;
`generate_esql` / `get_document_by_id` detours | **Follow-up PR** —
skill prompt tightening (baseline measures current behavior) |
| Trace evaluators returned `n/a` | Agent Builder spans on local Scout
ES, not golden cluster | Suite overrides `traceEsClient` to local Scout
ES |
| EIS connector setup flaky | Dot-prefixed `inferenceId` rejected;
manual cleanup between runs | Documented as operator setup issue —
**not** fixed in this PR (see out of scope) |
| Quality-only scoring hides waste | Rubric = 1 while trajectory = 0.30
and CostPerAlert = 0 | Added `ForbiddenTools`, `CostPerAlert`, strict
`Trajectory` evaluators |
| Only happy paths tested | Two golden cases insufficient for regression
signal | Added non-golden golden-path cases; clean-profile quality
cohort kept **on-demand** (not weekly) |

---

## What changed

### New evals package


`x-pack/solutions/security/packages/kbn-evals-suite-attack-discovery-agent-builder`

- Deterministic fixtures with **relative timestamps** (static
`@timestamp` caused live-retrieval to fail once alerts fell outside the
"last six hours" window).
- Custom evaluators: `WorkflowEvidence`, `StrictTrajectory`,
`ResponseSkillInvocation`, `ForbiddenTools`, `CostPerAlert`,
`AttackDiscoveryBasic`, `Criteria`, `Rubric`.
- `chat_client.ts` uses natural routing only — no harness
`configuration_overrides`; `expectedSkills` feeds Skill Invoked scoring.
- Trajectory precision excludes `load_skill` (framework routing cost
under natural routing).

### Dataset coverage

| Path | Type | What it validates |
|------|------|-------------------|
| `provided-alerts` | Golden | Alerts attached; agent skips retrieval |
| `live-retrieval` | Golden | Agent retrieves via default ES\|QL query |
| `multiple-alert-sets` | Non-golden | Multiple alert attachments in one
request |
| `missing-alert-retrieval` | Non-golden | Live retrieval returns zero
alerts; agent stops cleanly |
| `status-only` | Non-golden | User asks for `execution_uuid` status; no
new generation |
| `clean-profile` (scenario registry) | On-demand | Four portable-seeder
chains; **quality** evals (not weekly gate) |

### Infrastructure

- Scout config set: `evals_attack_discovery_agent_builder` (AD 2.0
feature flags).
- Suite registered in `evals.suites.json`, `package.json`,
`tsconfig.base.json`.
- Weekly pipeline: `llm_evals.yml` step for
`attack-discovery-agent-builder` with `EVAL_GREP` (golden-path only).
Clean profile remains in-repo for on-demand runs.

### Product improvements driven by evals

- **None in this PR** — skill-side routing prose changes moved to the
follow-up PR so this merge establishes a reproducible **main-baseline**
scorecard.

### Intentionally out of scope

This PR stays focused on the **AD2 eval suite**, **skill-side routing
prose**, and **weekly/on-demand CI registration**. The following were
explored during development but **reverted** to avoid unrelated platform
diffs:

- **`platform.core.generate_esql` tool description** — prefer-a-skill
boundary clause is a separate Agent Builder platform change, not part of
this evals PR.
- **`evals_tracing` EIS `inferenceId` normalization** — connector-cache
dot-prefix is an operator/local-setup concern; no Scout config change
ships here.
- **`@kbn/evals` `getToolCallSteps` `params` export** — the suite reads
`load_skill` arguments via local `getToolCallStepsWithParams()` instead
of extending the shared helper.

---

## How to run

```bash
node scripts/evals run --suite attack-discovery-agent-builder --connector eis-anthropic-claude-5-sonnet
```

Focused development:

```bash
node scripts/evals run --suite attack-discovery-agent-builder --grep "golden provided-alerts" --repetitions 1
```

---

## Follow-up (not blocking this PR)

- Further reduce live-retrieval variance (currently 9–14 tool calls;
CostPerAlert threshold is 5/alert).
- Scout ES stack stability under consecutive heavy runs (infrastructure,
not eval logic).
- Rubric judge flake on golden encoded-powershell export (`judge_failed`
/ orphaned `tool_result` — quality evaluators still pass).

## Related

- Skill-dev plugin lessons captured in
[elastic/agent-builder-skill-dev-cursor-plugin#149](elastic/agent-builder-skill-dev-cursor-plugin#149)
(`agent-builder-eval-playbook` skill).
- Does **not** merge with the legacy direct-generation cohort
(`kbn-evals-suite-attack-discovery`).

### 5. Weekly-matrix multi-model comparison (baseline vs after)

Local batch runs on **2026-07-13** (M4, `feat/search-skill-boundary`
worktree). LLM-gateway models skipped (`AD2_SKIP_GATEWAY_MODELS=true`).

| Cohort | TEST_RUN_ID | Playwright result |
|--------|-------------|-------------------|
| baseline | `ad2-baseline-20260713-112322` | **6/6 models — 5/5 tests
each** (`upstream/main` skill + `generate_esql`) |
| after (initial) | `ad2-after-20260713-144634` | **3/6 complete**
(Sonnet, Opus, Flash); 3 infra-failed |
| after (retry) | `ad2-after-retry-20260713-153105` | **1/3 recovered**
(GPT-OSS 5/5); Gemini 3.1 Pro + GPT-5.4 still infra-failed |

#### Infrastructure fix applied before retry

Dead-stack recovery was steering Playwright at `:5620` because
`.scout/servers/local.json` was clobbered by competing Scout boots while
the batch worker stayed on `:15001`/`:15000`. Batch runner now:

1. **`sync_scout_local_json()`** — patches (or recreates) `local.json`
hosts before every suite dispatch and after every boot/reboot.
2. **`ensure_stack_alive()`** — uses full two-pass `boot_stack()` on
dead-stack detection (not single-pass `boot_stack_once`), so EIS
connectors are re-baked after ES wipe.

#### Playwright pass matrix (merged after cohort)

| Model | Baseline | After (merged) | Notes |
|-------|----------|----------------|-------|
| Claude 4.6 Sonnet | 5/5 | **5/5** | after initial run |
| Claude 4.6 Opus | 5/5 | **5/5** | after initial run |
| Gemini 3.0 Flash | 5/5 | **5/5** | after initial run |
| Gemini 3.1 Pro | 5/5 | **2/5 fail** (retry also 2/5) | Stack died on
`:15001` during converse; later tests hit missing `local.json` |
| GPT-5.4 | 5/5 | **0/5** | Connector race at `playwright.config.ts`
load (`Evaluation connector id … was not found`) — both initial + retry
|
| GPT-OSS-120b | 5/5 | **5/5** | Retry succeeded on `:15001`/`:15000`
after port-sync fix |

#### Evaluator Δ table (Overall means, log-parsed)

Scores extracted from Playwright summary tables in batch logs (ephemeral
Scout ES on `:15000` was torn down before ES-indexed export;
`generate-ad2-comparison-report.mjs` could not query
`.evaluation-scores*` post-teardown). **Δ = after − baseline.**

| Model | PW | AdToolResult | WorkflowEvidence | trajectory |
ForbiddenTools | AttackDiscoveryBasic | Criteria | Rubric | Skill
Invoked | Tool Calls | Latency |

|-------|-----|-------------:|-----------------:|-----------:|---------------:|---------------------:|---------:|-------:|--------------:|-----------:|--------:|
| Claude 4.6 Sonnet | 5/5→5/5 | 0.50→0.50 (0) | 0.83→0.67 (−0.16) | 1→1
(0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 0.50→0.52 (+0.02) |
6.7→6.3 | 34.2s→31.4s |
| Claude 4.6 Opus | 5/5→5/5 | 0.83→0.67 (−0.16) | 0.50→0.83 (+0.33) |
0.80→0.75 (−0.05) | 0.33→1 (+0.67) | 0.50→0.67 (+0.17) | 1→1 (0) |
0.33→0.83 (+0.50) | 0.24→0.37 (+0.13) | 18.3→12.0 | 60.5s→42.2s |
| Gemini 3.0 Flash | 5/5→5/5 | 0.50→0.50 (0) | 0.50→0.50 (0) | 0.67→0.33
(−0.34) | 0.93→0.60 (−0.33) | 1→0.67 (−0.33) | 1→1 (0) | 1→0.67 (−0.33)
| 0.16→0.23 (+0.07) | 15.2→21.2 | 13.2s→19.6s |
| Gemini 3.1 Pro | 5/5→**2/5 fail** | — | — | — | — | — | — | — | — | —
| — |
| GPT-5.4 | 5/5→**0/5** | — | — | — | — | — | — | — | — | — | — |
| GPT-OSS-120b | 5/5→5/5 | 0.50→0.00 (−0.50) | 0.17→0.17 (0) | 1→— |
0.33→0.40 (+0.07) | 1→0.83 (−0.17) | 1→0.83 (−0.17) | 0.67→0.83 (+0.16)
| 0.43→0.00 (−0.43) | 5.2→12.2 | 52.9s→30.4s |

Baseline GPT-OSS row from `ad2-baseline-20260713-112322` run log; after
GPT-OSS row from retry run `ad2-after-retry-20260713-153105`.

#### Regression watch (completed models only)

On the **4 models with full 5/5 after Playwright completion** (Sonnet,
Opus, Flash, GPT-OSS):

- **No Criteria / AttackDiscoveryBasic regressions** — Criteria stays at
1.0 where scored; Basic dips on Flash (−0.33) and GPT-OSS (−0.17) but
remains ≥0.67.
- **Efficiency mixed:** Opus improves (fewer tool calls, lower latency);
Flash regresses on trajectory (−0.34) and tool calls (+6).
- **GPT-OSS quality metrics look lower** (AdToolResult 0, Skill Invoked
0) but this matches baseline's already-low AdTool/Skill scores — not
introduced by after-cohort hardening.

**Cannot conclude** for Gemini 3.1 Pro or GPT-5.4 — failures are
stack/connector infrastructure, not scored quality regressions.

#### Harness notes

- Stop competing Scout stacks before batch: `node scripts/evals stop`
- `TEST_KIBANA_PORT`/`TEST_ES_PORT` patch in `base.config.ts` required
when branch lacks `fix/weekly-evals-matrix` merge
- Batch runner port-sync + full `boot_stack()` recovery landed in
`run-security-evals-batch.sh` (skill-dev plugin, local)
- Remaining gap: per-model connector-registration wait before
`playwright.config.ts` loads (GPT-5.4 race)

### 6. Natural-routing convergence + OTLP trace evidence
(`ad2-full-natural-20260713-191333`)

**Harness change:** removed `configuration_overrides` /
`AD2_EVAL_FORCE_SKILL_ROUTING` — default Agent Builder router only;
`expectedSkills` is scoring-only.

**Run:** `ad2-full-natural-20260713-191333` on
`feat/search-skill-boundary` worktree (Claude 4.6 Sonnet, local Scout
`:15001`/`:15000`). **9/9 Playwright passed** (25.3m).

#### Representative trace cases (OTLP on local Scout ES — attach before
stack restart)

| Case | trace.id | Spans | Skill Invoked | ForbiddenTools | trajectory
| Tool Calls | Quality |

|------|----------|-------|---------------|----------------|------------|------------|---------|
| **live-retrieval** | `bf1f541ffa9a8e91af088dd526a74ae6` | 13 | **1.0**
| **1.0** | 0.75 | 4 | Criteria/Rubric/AdToolResult **1.0** |
| **clean-profile (encoded-powershell)** |
`f49eca1730170cf49fbf48b6a89fca0e` | 9 | **1.0** | **1.0** | 0.50 | 2 |
Criteria/Rubric/AdToolResult **1.0** |

**Live-retrieval OTLP tool sequence:** `load_skill` →
`get_default_esql_query` (custom) → `platform.core.execute_esql` →
`security.attack-discovery.run` (custom). No `generate_esql`, no
`attack-discovery-alert-retrieval-builder`.

**Clean-profile OTLP tool sequence:** `load_skill` →
`security.attack-discovery.run` with provided alert IDs.

**Skill-side fix (routing):** `attack_discovery_generator_skill.ts` —
routing keywords, `ALERT_RETRIEVAL_BOUNDARIES` (minimal
`get_default_esql_query` → `execute_esql` path), `STATUS_GUIDE` for
status-only queries.

**Local evidence artifacts** (operator worktree, gitignored):
`.ad2-pr-evidence/ad2-full-natural-20260713-191333/` — `README.md`,
`scores/score-summary.json`, `traces/*-otlp-projected.json`, raw OTLP
JSON.

### Golden-cluster OTLP trace IDs (exported 2026-07-13)

| Case | Golden `trace.id` | Run ID | Skill Invoked | ForbiddenTools |
Waterfall |

|------|-------------------|--------|---------------|----------------|-----------|
| **live-retrieval** (happy path) | `7c29a6966e8766ffe343c9e3f8516aeb` |
`ad2-golden-export-live-20260713-204938` | 1.0 | 1.0 | [APM
trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=7c29a6966e8766ffe343c9e3f8516aeb)
|
| **clean-profile encoded-powershell** |
`594f040b9f47df82f955107e6aaf5019` | `ad2-golden-export-20260713-203039`
| 1.0 | 1.0 | [APM
trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=594f040b9f47df82f955107e6aaf5019)
|

Local Scout OTLP snapshots (pre-teardown):
`bf1f541ffa9a8e91af088dd526a74ae6`, `f49eca1730170cf49fbf48b6a89fca0e` —
see `.ad2-pr-evidence/ad2-full-natural-20260713-191333/`.

### 7. Post-merge natural routing (2026-07-13)

Latest commit on this branch drops harness forcing, adds clean-profile
scenario-registry evals, and tunes trajectory precision to exclude
`load_skill` under natural routing.
### 8. Skill-baseline golden-path scorecard (2026-07-15)

**CI:** [Buildkite
#469078](https://buildkite.com/elastic/kibana-pull-request/builds/469078)
**PASS** on `208d3c253b0a` (skill prose reverted to upstream/main).

**Local run:** `feat/attack-discovery-agent-builder-evals` @
`208d3c253b0a`, EIS Claude 4.6 Sonnet, `--grep "golden |non-golden"`,
rep=1, Scout `evals_attack_discovery_agent_builder`. **5/5 Playwright
passed** (11.7m).

| Dataset | AdToolResult | Basic | Criteria | ForbiddenTools | Rubric |
Skill Invoked | Tool Calls | WorkflowEvidence | Trajectory |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| provided-alerts | 1 | 1 | 1 | 1 | 1 | 1 | 4 | 1 | 0.33 |
| live-retrieval (n=2) | 0 | 0.50 | **0** | **0** | 0.50 | 1 | **11.5**
| — | **0.20** |
| multiple-alert-sets | 0 | 1 | **0** | 1 | **0** | 1 | **9** | **0** |
0.14 |
| missing-alert-retrieval | 0 | 1 | — | 0 | 1 | 1 | 9 | — | 0.29 |
| status-only | 0 | 0 | — | 1 | 1 | 1 | 2 | — | 1 |

**Interpretation:** With **main skill prose**, routing works (Skill
Invoked = 1 everywhere) but **live-retrieval** and
**multiple-alert-sets** show efficiency/quality gaps — extra tool calls,
low Criteria/Rubric, missing WorkflowEvidence. These are the comparison
targets for skill hardening in follow-up
[#13](#13)
(not in this PR).

**Follow-up A/B (skill overlay from #13, same harness):** live-retrieval
Criteria 0→1, Rubric 0.50→1, ForbiddenTools 0→1, tool calls 11.5→4,
trajectory 0.20→0.83; multiple-alert-sets Criteria 0→1, Rubric 0→1,
WorkflowEvidence 0→1, tool calls 9→2. See #13 for full Δ table.

---------

Signed-off-by: Patryk Kopycinski <patryk.kopycinski@elastic.co>
Co-authored-by: kibanamachine <42973632+kibanamachine@users.noreply.github.com>
patrykkopycinski added a commit that referenced this pull request Aug 18, 2026
…277625)

## Summary

This PR introduces a dedicated, isolated evaluation suite for the
**Attack Discovery 2.0 Agent Builder** integration. It is intentionally
separate from the legacy direct-generation cohort so that we can measure
workflow, routing, tool efficiency, and output quality independently.

**The point of this PR is to show what focused evals buy you:** we did
not just add tests — running them surfaced routing noise, wasted tool
calls, and trace routing confusion, and those findings drove the **skill
and harness** changes included here.

**Baseline scope:** This PR ships the **eval suite + weekly CI wiring
only**. `attack-discovery-generator` skill prose is intentionally left
at **upstream/main** so follow-up PRs can measure improvement against
this baseline.

Committed evals use the **default Agent Builder router** — no
`configuration_overrides`. Dataset `expectedSkills` is scoring-only
(Skill Invoked evaluator).

---

## Eval profiles and CI cadence

This suite ships **two automated cohorts** in one package; they answer
different questions and run on different cadences.

| Profile | Spec | Seed | CI cadence | Primary question |
| --- | --- | --- | --- | --- |
| **Golden-path** | `attack_discovery_agent_builder.spec.ts` |
`fixtures.ts` (2 marker alerts) | **Weekly** — `llm_evals.yml` sets
`EVAL_GREP: 'golden |non-golden'` | Default-agent **routing**,
**tool/workflow plumbing**, efficiency gates |
| **Clean profile** | `clean_profile_provided_alerts.spec.ts` |
`scenario_registry/` (4 chains, 16 alerts + raw events) | **On-demand**
— run with `--grep "clean profile"` or full suite locally/Buildkite
on-demand | **Discovery quality** on realistic multi-stage chains
(portable-seeder `clean` parity) |
| **Full profile** | — | `ad-2.0-portable-seeder.py --profile full` |
**Not automated** in this PR | Signal vs **noise** (~150+ distractor
alerts); needs FPR/discrimination evaluators first |

**Do not merge golden-path and clean profile for weekly CI** —
golden-path includes live-retrieval, missing-retrieval, and status-only
cases that depend on the minimal marker fixture; clean profile is
provided-alerts-only on richer chain data.



## Evidence collected

### 1. Baseline (pre-prompt-tightening, first AD2 golden runs)

| Path | Tool calls | Input tokens | Notes |
|------|-----------|--------------|-------|
| **provided-alerts** | 5 | ~266k | Skill invoked, AD tool ran |
| **live-retrieval** | 13 | ~721k | Extra retrieval/corroboration tools
|

**Problem:** live-retrieval was doing 2.6× the tool calls and 2.7× the
tokens for the same 2-alert fixture. Quality was fine, but efficiency
was not.

### 2. After prompt tightening + new evaluators (clean stack restart)

| Path | Tool calls | Input tokens | AdToolResult | Basic | Criteria |
Rubric | Skill Invoked | Workflow | Trajectory |

|------|-----------|--------------|--------------|-------|----------|--------|---------------|----------|------------|
| **provided-alerts** | **4** | ~166k | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| **live-retrieval** | **10** | ~387k | 1 | 1 | 1 | 1 | 1 | 1 | ~0.30 |

**Improvement vs baseline:**

| Metric | provided-alerts | live-retrieval |
|--------|-----------------|----------------|
| Tool calls | 5 → **4** (−20%) | 13 → **10** (−23%) |
| Input tokens | ~266k → **~166k** (−38%) | ~721k → **~387k** (−46%) |
| Quality evaluators | already 1.0 | already 1.0 |

Trajectory on live-retrieval is intentionally low (~0.30) because the
agent legitimately runs `get_default_esql_query` → `execute_esql` →
`run`, but still adds corroboration detours. The new `ForbiddenTools`
and `CostPerAlert` evaluators make those detours visible instead of
hiding behind a quality score of 1.

### 3. Live-retrieval variance study (3 consecutive repetitions)

Even after tightening, live-retrieval is **non-deterministic**:

| Rep | Tool calls | Input tokens | Latency | ForbiddenTools |
CostPerAlert | Quality |

|-----|-----------|--------------|---------|----------------|--------------|---------|
| 1 | 9 | 413k | 37s | 0 | 1 | 1 |
| 2 | 14 | 657k | 135s | 1 | 0 | 1 |
| 3 | 11 | 586k | 38s | 0 | 0 | 1 |
| **mean** | **11.3** | **552k** | **70s** | **0.33** | **0.67** |
**1.0** |

- **ForbiddenTools** still leaks occasionally
(`platform.core.get_document_by_id`, `platform.core.generate_esql`).
- **CostPerAlert** fails 2/3 runs (threshold: 5 calls/alert × 2 alerts =
10; rep 2 hit 14).
- **Quality stays at 1.0** across all reps — efficiency is the remaining
gap, not output correctness.

### 4. What the evals revealed (and what we changed because of it)

| Finding | Evidence | Change in this PR |
|---------|----------|-------------------|
| Default router sometimes picks wrong skill | Live-retrieval misroutes
(baseline skill prose) | **Follow-up PR** — skill hardening moved out so
this PR stays baseline |
| Speculative tool calls inflate cost | 9–14 calls for 2 alerts;
`generate_esql` / `get_document_by_id` detours | **Follow-up PR** —
skill prompt tightening (baseline measures current behavior) |
| Trace evaluators returned `n/a` | Agent Builder spans on local Scout
ES, not golden cluster | Suite overrides `traceEsClient` to local Scout
ES |
| EIS connector setup flaky | Dot-prefixed `inferenceId` rejected;
manual cleanup between runs | Documented as operator setup issue —
**not** fixed in this PR (see out of scope) |
| Quality-only scoring hides waste | Rubric = 1 while trajectory = 0.30
and CostPerAlert = 0 | Added `ForbiddenTools`, `CostPerAlert`, strict
`Trajectory` evaluators |
| Only happy paths tested | Two golden cases insufficient for regression
signal | Added non-golden golden-path cases; clean-profile quality
cohort kept **on-demand** (not weekly) |

---

## What changed

### New evals package


`x-pack/solutions/security/packages/kbn-evals-suite-attack-discovery-agent-builder`

- Deterministic fixtures with **relative timestamps** (static
`@timestamp` caused live-retrieval to fail once alerts fell outside the
"last six hours" window).
- Custom evaluators: `WorkflowEvidence`, `StrictTrajectory`,
`ResponseSkillInvocation`, `ForbiddenTools`, `CostPerAlert`,
`AttackDiscoveryBasic`, `Criteria`, `Rubric`.
- `chat_client.ts` uses natural routing only — no harness
`configuration_overrides`; `expectedSkills` feeds Skill Invoked scoring.
- Trajectory precision excludes `load_skill` (framework routing cost
under natural routing).

### Dataset coverage

| Path | Type | What it validates |
|------|------|-------------------|
| `provided-alerts` | Golden | Alerts attached; agent skips retrieval |
| `live-retrieval` | Golden | Agent retrieves via default ES\|QL query |
| `multiple-alert-sets` | Non-golden | Multiple alert attachments in one
request |
| `missing-alert-retrieval` | Non-golden | Live retrieval returns zero
alerts; agent stops cleanly |
| `status-only` | Non-golden | User asks for `execution_uuid` status; no
new generation |
| `clean-profile` (scenario registry) | On-demand | Four portable-seeder
chains; **quality** evals (not weekly gate) |

### Infrastructure

- Scout config set: `evals_attack_discovery_agent_builder` (AD 2.0
feature flags).
- Suite registered in `evals.suites.json`, `package.json`,
`tsconfig.base.json`.
- Weekly pipeline: `llm_evals.yml` step for
`attack-discovery-agent-builder` with `EVAL_GREP` (golden-path only).
Clean profile remains in-repo for on-demand runs.

### Product improvements driven by evals

- **None in this PR** — skill-side routing prose changes moved to the
follow-up PR so this merge establishes a reproducible **main-baseline**
scorecard.

### Intentionally out of scope

This PR stays focused on the **AD2 eval suite**, **skill-side routing
prose**, and **weekly/on-demand CI registration**. The following were
explored during development but **reverted** to avoid unrelated platform
diffs:

- **`platform.core.generate_esql` tool description** — prefer-a-skill
boundary clause is a separate Agent Builder platform change, not part of
this evals PR.
- **`evals_tracing` EIS `inferenceId` normalization** — connector-cache
dot-prefix is an operator/local-setup concern; no Scout config change
ships here.
- **`@kbn/evals` `getToolCallSteps` `params` export** — the suite reads
`load_skill` arguments via local `getToolCallStepsWithParams()` instead
of extending the shared helper.

---

## How to run

```bash
node scripts/evals run --suite attack-discovery-agent-builder --connector eis-anthropic-claude-5-sonnet
```

Focused development:

```bash
node scripts/evals run --suite attack-discovery-agent-builder --grep "golden provided-alerts" --repetitions 1
```

---

## Follow-up (not blocking this PR)

- Further reduce live-retrieval variance (currently 9–14 tool calls;
CostPerAlert threshold is 5/alert).
- Scout ES stack stability under consecutive heavy runs (infrastructure,
not eval logic).
- Rubric judge flake on golden encoded-powershell export (`judge_failed`
/ orphaned `tool_result` — quality evaluators still pass).

## Related

- Skill-dev plugin lessons captured in
[elastic/agent-builder-skill-dev-cursor-plugin#149](elastic/agent-builder-skill-dev-cursor-plugin#149)
(`agent-builder-eval-playbook` skill).
- Does **not** merge with the legacy direct-generation cohort
(`kbn-evals-suite-attack-discovery`).

### 5. Weekly-matrix multi-model comparison (baseline vs after)

Local batch runs on **2026-07-13** (M4, `feat/search-skill-boundary`
worktree). LLM-gateway models skipped (`AD2_SKIP_GATEWAY_MODELS=true`).

| Cohort | TEST_RUN_ID | Playwright result |
|--------|-------------|-------------------|
| baseline | `ad2-baseline-20260713-112322` | **6/6 models — 5/5 tests
each** (`upstream/main` skill + `generate_esql`) |
| after (initial) | `ad2-after-20260713-144634` | **3/6 complete**
(Sonnet, Opus, Flash); 3 infra-failed |
| after (retry) | `ad2-after-retry-20260713-153105` | **1/3 recovered**
(GPT-OSS 5/5); Gemini 3.1 Pro + GPT-5.4 still infra-failed |

#### Infrastructure fix applied before retry

Dead-stack recovery was steering Playwright at `:5620` because
`.scout/servers/local.json` was clobbered by competing Scout boots while
the batch worker stayed on `:15001`/`:15000`. Batch runner now:

1. **`sync_scout_local_json()`** — patches (or recreates) `local.json`
hosts before every suite dispatch and after every boot/reboot.
2. **`ensure_stack_alive()`** — uses full two-pass `boot_stack()` on
dead-stack detection (not single-pass `boot_stack_once`), so EIS
connectors are re-baked after ES wipe.

#### Playwright pass matrix (merged after cohort)

| Model | Baseline | After (merged) | Notes |
|-------|----------|----------------|-------|
| Claude 4.6 Sonnet | 5/5 | **5/5** | after initial run |
| Claude 4.6 Opus | 5/5 | **5/5** | after initial run |
| Gemini 3.0 Flash | 5/5 | **5/5** | after initial run |
| Gemini 3.1 Pro | 5/5 | **2/5 fail** (retry also 2/5) | Stack died on
`:15001` during converse; later tests hit missing `local.json` |
| GPT-5.4 | 5/5 | **0/5** | Connector race at `playwright.config.ts`
load (`Evaluation connector id … was not found`) — both initial + retry
|
| GPT-OSS-120b | 5/5 | **5/5** | Retry succeeded on `:15001`/`:15000`
after port-sync fix |

#### Evaluator Δ table (Overall means, log-parsed)

Scores extracted from Playwright summary tables in batch logs (ephemeral
Scout ES on `:15000` was torn down before ES-indexed export;
`generate-ad2-comparison-report.mjs` could not query
`.evaluation-scores*` post-teardown). **Δ = after − baseline.**

| Model | PW | AdToolResult | WorkflowEvidence | trajectory |
ForbiddenTools | AttackDiscoveryBasic | Criteria | Rubric | Skill
Invoked | Tool Calls | Latency |

|-------|-----|-------------:|-----------------:|-----------:|---------------:|---------------------:|---------:|-------:|--------------:|-----------:|--------:|
| Claude 4.6 Sonnet | 5/5→5/5 | 0.50→0.50 (0) | 0.83→0.67 (−0.16) | 1→1
(0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 0.50→0.52 (+0.02) |
6.7→6.3 | 34.2s→31.4s |
| Claude 4.6 Opus | 5/5→5/5 | 0.83→0.67 (−0.16) | 0.50→0.83 (+0.33) |
0.80→0.75 (−0.05) | 0.33→1 (+0.67) | 0.50→0.67 (+0.17) | 1→1 (0) |
0.33→0.83 (+0.50) | 0.24→0.37 (+0.13) | 18.3→12.0 | 60.5s→42.2s |
| Gemini 3.0 Flash | 5/5→5/5 | 0.50→0.50 (0) | 0.50→0.50 (0) | 0.67→0.33
(−0.34) | 0.93→0.60 (−0.33) | 1→0.67 (−0.33) | 1→1 (0) | 1→0.67 (−0.33)
| 0.16→0.23 (+0.07) | 15.2→21.2 | 13.2s→19.6s |
| Gemini 3.1 Pro | 5/5→**2/5 fail** | — | — | — | — | — | — | — | — | —
| — |
| GPT-5.4 | 5/5→**0/5** | — | — | — | — | — | — | — | — | — | — |
| GPT-OSS-120b | 5/5→5/5 | 0.50→0.00 (−0.50) | 0.17→0.17 (0) | 1→— |
0.33→0.40 (+0.07) | 1→0.83 (−0.17) | 1→0.83 (−0.17) | 0.67→0.83 (+0.16)
| 0.43→0.00 (−0.43) | 5.2→12.2 | 52.9s→30.4s |

Baseline GPT-OSS row from `ad2-baseline-20260713-112322` run log; after
GPT-OSS row from retry run `ad2-after-retry-20260713-153105`.

#### Regression watch (completed models only)

On the **4 models with full 5/5 after Playwright completion** (Sonnet,
Opus, Flash, GPT-OSS):

- **No Criteria / AttackDiscoveryBasic regressions** — Criteria stays at
1.0 where scored; Basic dips on Flash (−0.33) and GPT-OSS (−0.17) but
remains ≥0.67.
- **Efficiency mixed:** Opus improves (fewer tool calls, lower latency);
Flash regresses on trajectory (−0.34) and tool calls (+6).
- **GPT-OSS quality metrics look lower** (AdToolResult 0, Skill Invoked
0) but this matches baseline's already-low AdTool/Skill scores — not
introduced by after-cohort hardening.

**Cannot conclude** for Gemini 3.1 Pro or GPT-5.4 — failures are
stack/connector infrastructure, not scored quality regressions.

#### Harness notes

- Stop competing Scout stacks before batch: `node scripts/evals stop`
- `TEST_KIBANA_PORT`/`TEST_ES_PORT` patch in `base.config.ts` required
when branch lacks `fix/weekly-evals-matrix` merge
- Batch runner port-sync + full `boot_stack()` recovery landed in
`run-security-evals-batch.sh` (skill-dev plugin, local)
- Remaining gap: per-model connector-registration wait before
`playwright.config.ts` loads (GPT-5.4 race)

### 6. Natural-routing convergence + OTLP trace evidence
(`ad2-full-natural-20260713-191333`)

**Harness change:** removed `configuration_overrides` /
`AD2_EVAL_FORCE_SKILL_ROUTING` — default Agent Builder router only;
`expectedSkills` is scoring-only.

**Run:** `ad2-full-natural-20260713-191333` on
`feat/search-skill-boundary` worktree (Claude 4.6 Sonnet, local Scout
`:15001`/`:15000`). **9/9 Playwright passed** (25.3m).

#### Representative trace cases (OTLP on local Scout ES — attach before
stack restart)

| Case | trace.id | Spans | Skill Invoked | ForbiddenTools | trajectory
| Tool Calls | Quality |

|------|----------|-------|---------------|----------------|------------|------------|---------|
| **live-retrieval** | `bf1f541ffa9a8e91af088dd526a74ae6` | 13 | **1.0**
| **1.0** | 0.75 | 4 | Criteria/Rubric/AdToolResult **1.0** |
| **clean-profile (encoded-powershell)** |
`f49eca1730170cf49fbf48b6a89fca0e` | 9 | **1.0** | **1.0** | 0.50 | 2 |
Criteria/Rubric/AdToolResult **1.0** |

**Live-retrieval OTLP tool sequence:** `load_skill` →
`get_default_esql_query` (custom) → `platform.core.execute_esql` →
`security.attack-discovery.run` (custom). No `generate_esql`, no
`attack-discovery-alert-retrieval-builder`.

**Clean-profile OTLP tool sequence:** `load_skill` →
`security.attack-discovery.run` with provided alert IDs.

**Skill-side fix (routing):** `attack_discovery_generator_skill.ts` —
routing keywords, `ALERT_RETRIEVAL_BOUNDARIES` (minimal
`get_default_esql_query` → `execute_esql` path), `STATUS_GUIDE` for
status-only queries.

**Local evidence artifacts** (operator worktree, gitignored):
`.ad2-pr-evidence/ad2-full-natural-20260713-191333/` — `README.md`,
`scores/score-summary.json`, `traces/*-otlp-projected.json`, raw OTLP
JSON.

### Golden-cluster OTLP trace IDs (exported 2026-07-13)

| Case | Golden `trace.id` | Run ID | Skill Invoked | ForbiddenTools |
Waterfall |

|------|-------------------|--------|---------------|----------------|-----------|
| **live-retrieval** (happy path) | `7c29a6966e8766ffe343c9e3f8516aeb` |
`ad2-golden-export-live-20260713-204938` | 1.0 | 1.0 | [APM
trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=7c29a6966e8766ffe343c9e3f8516aeb)
|
| **clean-profile encoded-powershell** |
`594f040b9f47df82f955107e6aaf5019` | `ad2-golden-export-20260713-203039`
| 1.0 | 1.0 | [APM
trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=594f040b9f47df82f955107e6aaf5019)
|

Local Scout OTLP snapshots (pre-teardown):
`bf1f541ffa9a8e91af088dd526a74ae6`, `f49eca1730170cf49fbf48b6a89fca0e` —
see `.ad2-pr-evidence/ad2-full-natural-20260713-191333/`.

### 7. Post-merge natural routing (2026-07-13)

Latest commit on this branch drops harness forcing, adds clean-profile
scenario-registry evals, and tunes trajectory precision to exclude
`load_skill` under natural routing.
### 8. Skill-baseline golden-path scorecard (2026-07-15)

**CI:** [Buildkite
#469078](https://buildkite.com/elastic/kibana-pull-request/builds/469078)
**PASS** on `208d3c253b0a` (skill prose reverted to upstream/main).

**Local run:** `feat/attack-discovery-agent-builder-evals` @
`208d3c253b0a`, EIS Claude 4.6 Sonnet, `--grep "golden |non-golden"`,
rep=1, Scout `evals_attack_discovery_agent_builder`. **5/5 Playwright
passed** (11.7m).

| Dataset | AdToolResult | Basic | Criteria | ForbiddenTools | Rubric |
Skill Invoked | Tool Calls | WorkflowEvidence | Trajectory |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| provided-alerts | 1 | 1 | 1 | 1 | 1 | 1 | 4 | 1 | 0.33 |
| live-retrieval (n=2) | 0 | 0.50 | **0** | **0** | 0.50 | 1 | **11.5**
| — | **0.20** |
| multiple-alert-sets | 0 | 1 | **0** | 1 | **0** | 1 | **9** | **0** |
0.14 |
| missing-alert-retrieval | 0 | 1 | — | 0 | 1 | 1 | 9 | — | 0.29 |
| status-only | 0 | 0 | — | 1 | 1 | 1 | 2 | — | 1 |

**Interpretation:** With **main skill prose**, routing works (Skill
Invoked = 1 everywhere) but **live-retrieval** and
**multiple-alert-sets** show efficiency/quality gaps — extra tool calls,
low Criteria/Rubric, missing WorkflowEvidence. These are the comparison
targets for skill hardening in follow-up
[#13](#13)
(not in this PR).

**Follow-up A/B (skill overlay from #13, same harness):** live-retrieval
Criteria 0→1, Rubric 0.50→1, ForbiddenTools 0→1, tool calls 11.5→4,
trajectory 0.20→0.83; multiple-alert-sets Criteria 0→1, Rubric 0→1,
WorkflowEvidence 0→1, tool calls 9→2. See #13 for full Δ table.

---------

Signed-off-by: Patryk Kopycinski <patryk.kopycinski@elastic.co>
Co-authored-by: kibanamachine <42973632+kibanamachine@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant