Repository navigation
[evals] AD2 full-profile on-demand discrimination (follow-up) - #13
Draft
patrykkopycinski wants to merge 3 commits into
Draft
patrykkopycinski wants to merge 3 commits into
patrykkopycinski wants to merge 3 commits into
Conversation
Extend the AD2 scenario registry with portable-seeder full profile seeding (seven signal chains plus deterministic background and Defender noise), add FPR evaluators (NoiseFalsePositive, DiscoveryCountCap, MinValidatedDiscovery), and register an on-demand live-retrieval spec — not wired to weekly CI.
Move routing prose improvements out of the baseline eval PR so this follow-up can be scored against elastic#277625 main-baseline skill behavior.
patrykkopycinski
force-pushed
the
feat/ad2-full-profile-evals
branch
from
July 14, 2026 22:00
f2bf371 to
d288ae7
Compare
Finding #7: Add trigger phrases to SKILL_DESCRIPTION so TF-IDF routing and weaker models match paraphrases like 'run across these alert sets'. Finding #5: Move Field Syntax before Mode C and Missed Detection Closure in SKILL_CONTENT so the core generate/status instructions appear first, giving smaller models more usable context window before the appendix-like post-report and gate sections. Finding #8: Prefer insights from the run tool's attack_discoveries field (parseInsightsFromToolResult) over regex-parsing the last ```json block from the message. Falls back to message parsing with backwards search for the insights key, fixing the case where a model emits a proposed ES|QL rule after the insights JSON block.
patrykkopycinski
pushed a commit
that referenced
this pull request
Aug 5, 2026
## Summary Set `connect.timeout = 60s` on the undici `Agent` used by `KbnClientRequester` (https path only). ## Why elastic#268531 migrated `KbnClient` from axios to native fetch but did not override undici's 10s `connect.timeout` default. Axios had no equivalent cutoff, so FTR callers talking to a busy local Kibana started failing once that PR landed. The `kibana-streams-performance` weekly pipeline went red in builds #9, #11, #12, and #13 with: ``` ConnectTimeoutError: Connect Timeout Error (attempted address: localhost:5620, timeout: 10000ms) ``` The `10000ms` is undici's default. Bisect: build #8 last green (2026-05-11) → #9 first red (2026-05-18), with elastic#268531 in the window. ## What changed `src/platform/packages/shared/kbn-kbn-client/src/kbn_client/kbn_client_requester.ts`: one constant, one option on the https `Agent`. http branch unchanged. ## Related Regression introduced in elastic#268531. Companion streams perf PR: elastic#270636. ## Validation https://buildkite.com/elastic/kibana-streams-performance/builds/14
patrykkopycinski
added a commit
to elastic/kibana
that referenced
this pull request
Aug 11, 2026
## Summary This PR introduces a dedicated, isolated evaluation suite for the **Attack Discovery 2.0 Agent Builder** integration. It is intentionally separate from the legacy direct-generation cohort so that we can measure workflow, routing, tool efficiency, and output quality independently. **The point of this PR is to show what focused evals buy you:** we did not just add tests — running them surfaced routing noise, wasted tool calls, and trace routing confusion, and those findings drove the **skill and harness** changes included here. **Baseline scope:** This PR ships the **eval suite + weekly CI wiring only**. `attack-discovery-generator` skill prose is intentionally left at **upstream/main** so follow-up PRs can measure improvement against this baseline. Committed evals use the **default Agent Builder router** — no `configuration_overrides`. Dataset `expectedSkills` is scoring-only (Skill Invoked evaluator). --- ## Eval profiles and CI cadence This suite ships **two automated cohorts** in one package; they answer different questions and run on different cadences. | Profile | Spec | Seed | CI cadence | Primary question | | --- | --- | --- | --- | --- | | **Golden-path** | `attack_discovery_agent_builder.spec.ts` | `fixtures.ts` (2 marker alerts) | **Weekly** — `llm_evals.yml` sets `EVAL_GREP: 'golden |non-golden'` | Default-agent **routing**, **tool/workflow plumbing**, efficiency gates | | **Clean profile** | `clean_profile_provided_alerts.spec.ts` | `scenario_registry/` (4 chains, 16 alerts + raw events) | **On-demand** — run with `--grep "clean profile"` or full suite locally/Buildkite on-demand | **Discovery quality** on realistic multi-stage chains (portable-seeder `clean` parity) | | **Full profile** | — | `ad-2.0-portable-seeder.py --profile full` | **Not automated** in this PR | Signal vs **noise** (~150+ distractor alerts); needs FPR/discrimination evaluators first | **Do not merge golden-path and clean profile for weekly CI** — golden-path includes live-retrieval, missing-retrieval, and status-only cases that depend on the minimal marker fixture; clean profile is provided-alerts-only on richer chain data. ## Evidence collected ### 1. Baseline (pre-prompt-tightening, first AD2 golden runs) | Path | Tool calls | Input tokens | Notes | |------|-----------|--------------|-------| | **provided-alerts** | 5 | ~266k | Skill invoked, AD tool ran | | **live-retrieval** | 13 | ~721k | Extra retrieval/corroboration tools | **Problem:** live-retrieval was doing 2.6× the tool calls and 2.7× the tokens for the same 2-alert fixture. Quality was fine, but efficiency was not. ### 2. After prompt tightening + new evaluators (clean stack restart) | Path | Tool calls | Input tokens | AdToolResult | Basic | Criteria | Rubric | Skill Invoked | Workflow | Trajectory | |------|-----------|--------------|--------------|-------|----------|--------|---------------|----------|------------| | **provided-alerts** | **4** | ~166k | 1 | 1 | 1 | 1 | 1 | 1 | 1 | | **live-retrieval** | **10** | ~387k | 1 | 1 | 1 | 1 | 1 | 1 | ~0.30 | **Improvement vs baseline:** | Metric | provided-alerts | live-retrieval | |--------|-----------------|----------------| | Tool calls | 5 → **4** (−20%) | 13 → **10** (−23%) | | Input tokens | ~266k → **~166k** (−38%) | ~721k → **~387k** (−46%) | | Quality evaluators | already 1.0 | already 1.0 | Trajectory on live-retrieval is intentionally low (~0.30) because the agent legitimately runs `get_default_esql_query` → `execute_esql` → `run`, but still adds corroboration detours. The new `ForbiddenTools` and `CostPerAlert` evaluators make those detours visible instead of hiding behind a quality score of 1. ### 3. Live-retrieval variance study (3 consecutive repetitions) Even after tightening, live-retrieval is **non-deterministic**: | Rep | Tool calls | Input tokens | Latency | ForbiddenTools | CostPerAlert | Quality | |-----|-----------|--------------|---------|----------------|--------------|---------| | 1 | 9 | 413k | 37s | 0 | 1 | 1 | | 2 | 14 | 657k | 135s | 1 | 0 | 1 | | 3 | 11 | 586k | 38s | 0 | 0 | 1 | | **mean** | **11.3** | **552k** | **70s** | **0.33** | **0.67** | **1.0** | - **ForbiddenTools** still leaks occasionally (`platform.core.get_document_by_id`, `platform.core.generate_esql`). - **CostPerAlert** fails 2/3 runs (threshold: 5 calls/alert × 2 alerts = 10; rep 2 hit 14). - **Quality stays at 1.0** across all reps — efficiency is the remaining gap, not output correctness. ### 4. What the evals revealed (and what we changed because of it) | Finding | Evidence | Change in this PR | |---------|----------|-------------------| | Default router sometimes picks wrong skill | Live-retrieval misroutes (baseline skill prose) | **Follow-up PR** — skill hardening moved out so this PR stays baseline | | Speculative tool calls inflate cost | 9–14 calls for 2 alerts; `generate_esql` / `get_document_by_id` detours | **Follow-up PR** — skill prompt tightening (baseline measures current behavior) | | Trace evaluators returned `n/a` | Agent Builder spans on local Scout ES, not golden cluster | Suite overrides `traceEsClient` to local Scout ES | | EIS connector setup flaky | Dot-prefixed `inferenceId` rejected; manual cleanup between runs | Documented as operator setup issue — **not** fixed in this PR (see out of scope) | | Quality-only scoring hides waste | Rubric = 1 while trajectory = 0.30 and CostPerAlert = 0 | Added `ForbiddenTools`, `CostPerAlert`, strict `Trajectory` evaluators | | Only happy paths tested | Two golden cases insufficient for regression signal | Added non-golden golden-path cases; clean-profile quality cohort kept **on-demand** (not weekly) | --- ## What changed ### New evals package `x-pack/solutions/security/packages/kbn-evals-suite-attack-discovery-agent-builder` - Deterministic fixtures with **relative timestamps** (static `@timestamp` caused live-retrieval to fail once alerts fell outside the "last six hours" window). - Custom evaluators: `WorkflowEvidence`, `StrictTrajectory`, `ResponseSkillInvocation`, `ForbiddenTools`, `CostPerAlert`, `AttackDiscoveryBasic`, `Criteria`, `Rubric`. - `chat_client.ts` uses natural routing only — no harness `configuration_overrides`; `expectedSkills` feeds Skill Invoked scoring. - Trajectory precision excludes `load_skill` (framework routing cost under natural routing). ### Dataset coverage | Path | Type | What it validates | |------|------|-------------------| | `provided-alerts` | Golden | Alerts attached; agent skips retrieval | | `live-retrieval` | Golden | Agent retrieves via default ES\|QL query | | `multiple-alert-sets` | Non-golden | Multiple alert attachments in one request | | `missing-alert-retrieval` | Non-golden | Live retrieval returns zero alerts; agent stops cleanly | | `status-only` | Non-golden | User asks for `execution_uuid` status; no new generation | | `clean-profile` (scenario registry) | On-demand | Four portable-seeder chains; **quality** evals (not weekly gate) | ### Infrastructure - Scout config set: `evals_attack_discovery_agent_builder` (AD 2.0 feature flags). - Suite registered in `evals.suites.json`, `package.json`, `tsconfig.base.json`. - Weekly pipeline: `llm_evals.yml` step for `attack-discovery-agent-builder` with `EVAL_GREP` (golden-path only). Clean profile remains in-repo for on-demand runs. ### Product improvements driven by evals - **None in this PR** — skill-side routing prose changes moved to the follow-up PR so this merge establishes a reproducible **main-baseline** scorecard. ### Intentionally out of scope This PR stays focused on the **AD2 eval suite**, **skill-side routing prose**, and **weekly/on-demand CI registration**. The following were explored during development but **reverted** to avoid unrelated platform diffs: - **`platform.core.generate_esql` tool description** — prefer-a-skill boundary clause is a separate Agent Builder platform change, not part of this evals PR. - **`evals_tracing` EIS `inferenceId` normalization** — connector-cache dot-prefix is an operator/local-setup concern; no Scout config change ships here. - **`@kbn/evals` `getToolCallSteps` `params` export** — the suite reads `load_skill` arguments via local `getToolCallStepsWithParams()` instead of extending the shared helper. --- ## How to run ```bash node scripts/evals run --suite attack-discovery-agent-builder --connector eis-anthropic-claude-5-sonnet ``` Focused development: ```bash node scripts/evals run --suite attack-discovery-agent-builder --grep "golden provided-alerts" --repetitions 1 ``` --- ## Follow-up (not blocking this PR) - Further reduce live-retrieval variance (currently 9–14 tool calls; CostPerAlert threshold is 5/alert). - Scout ES stack stability under consecutive heavy runs (infrastructure, not eval logic). - Rubric judge flake on golden encoded-powershell export (`judge_failed` / orphaned `tool_result` — quality evaluators still pass). ## Related - Skill-dev plugin lessons captured in [elastic/agent-builder-skill-dev-cursor-plugin#149](elastic/agent-builder-skill-dev-cursor-plugin#149) (`agent-builder-eval-playbook` skill). - Does **not** merge with the legacy direct-generation cohort (`kbn-evals-suite-attack-discovery`). ### 5. Weekly-matrix multi-model comparison (baseline vs after) Local batch runs on **2026-07-13** (M4, `feat/search-skill-boundary` worktree). LLM-gateway models skipped (`AD2_SKIP_GATEWAY_MODELS=true`). | Cohort | TEST_RUN_ID | Playwright result | |--------|-------------|-------------------| | baseline | `ad2-baseline-20260713-112322` | **6/6 models — 5/5 tests each** (`upstream/main` skill + `generate_esql`) | | after (initial) | `ad2-after-20260713-144634` | **3/6 complete** (Sonnet, Opus, Flash); 3 infra-failed | | after (retry) | `ad2-after-retry-20260713-153105` | **1/3 recovered** (GPT-OSS 5/5); Gemini 3.1 Pro + GPT-5.4 still infra-failed | #### Infrastructure fix applied before retry Dead-stack recovery was steering Playwright at `:5620` because `.scout/servers/local.json` was clobbered by competing Scout boots while the batch worker stayed on `:15001`/`:15000`. Batch runner now: 1. **`sync_scout_local_json()`** — patches (or recreates) `local.json` hosts before every suite dispatch and after every boot/reboot. 2. **`ensure_stack_alive()`** — uses full two-pass `boot_stack()` on dead-stack detection (not single-pass `boot_stack_once`), so EIS connectors are re-baked after ES wipe. #### Playwright pass matrix (merged after cohort) | Model | Baseline | After (merged) | Notes | |-------|----------|----------------|-------| | Claude 4.6 Sonnet | 5/5 | **5/5** | after initial run | | Claude 4.6 Opus | 5/5 | **5/5** | after initial run | | Gemini 3.0 Flash | 5/5 | **5/5** | after initial run | | Gemini 3.1 Pro | 5/5 | **2/5 fail** (retry also 2/5) | Stack died on `:15001` during converse; later tests hit missing `local.json` | | GPT-5.4 | 5/5 | **0/5** | Connector race at `playwright.config.ts` load (`Evaluation connector id … was not found`) — both initial + retry | | GPT-OSS-120b | 5/5 | **5/5** | Retry succeeded on `:15001`/`:15000` after port-sync fix | #### Evaluator Δ table (Overall means, log-parsed) Scores extracted from Playwright summary tables in batch logs (ephemeral Scout ES on `:15000` was torn down before ES-indexed export; `generate-ad2-comparison-report.mjs` could not query `.evaluation-scores*` post-teardown). **Δ = after − baseline.** | Model | PW | AdToolResult | WorkflowEvidence | trajectory | ForbiddenTools | AttackDiscoveryBasic | Criteria | Rubric | Skill Invoked | Tool Calls | Latency | |-------|-----|-------------:|-----------------:|-----------:|---------------:|---------------------:|---------:|-------:|--------------:|-----------:|--------:| | Claude 4.6 Sonnet | 5/5→5/5 | 0.50→0.50 (0) | 0.83→0.67 (−0.16) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 0.50→0.52 (+0.02) | 6.7→6.3 | 34.2s→31.4s | | Claude 4.6 Opus | 5/5→5/5 | 0.83→0.67 (−0.16) | 0.50→0.83 (+0.33) | 0.80→0.75 (−0.05) | 0.33→1 (+0.67) | 0.50→0.67 (+0.17) | 1→1 (0) | 0.33→0.83 (+0.50) | 0.24→0.37 (+0.13) | 18.3→12.0 | 60.5s→42.2s | | Gemini 3.0 Flash | 5/5→5/5 | 0.50→0.50 (0) | 0.50→0.50 (0) | 0.67→0.33 (−0.34) | 0.93→0.60 (−0.33) | 1→0.67 (−0.33) | 1→1 (0) | 1→0.67 (−0.33) | 0.16→0.23 (+0.07) | 15.2→21.2 | 13.2s→19.6s | | Gemini 3.1 Pro | 5/5→**2/5 fail** | — | — | — | — | — | — | — | — | — | — | | GPT-5.4 | 5/5→**0/5** | — | — | — | — | — | — | — | — | — | — | | GPT-OSS-120b | 5/5→5/5 | 0.50→0.00 (−0.50) | 0.17→0.17 (0) | 1→— | 0.33→0.40 (+0.07) | 1→0.83 (−0.17) | 1→0.83 (−0.17) | 0.67→0.83 (+0.16) | 0.43→0.00 (−0.43) | 5.2→12.2 | 52.9s→30.4s | Baseline GPT-OSS row from `ad2-baseline-20260713-112322` run log; after GPT-OSS row from retry run `ad2-after-retry-20260713-153105`. #### Regression watch (completed models only) On the **4 models with full 5/5 after Playwright completion** (Sonnet, Opus, Flash, GPT-OSS): - **No Criteria / AttackDiscoveryBasic regressions** — Criteria stays at 1.0 where scored; Basic dips on Flash (−0.33) and GPT-OSS (−0.17) but remains ≥0.67. - **Efficiency mixed:** Opus improves (fewer tool calls, lower latency); Flash regresses on trajectory (−0.34) and tool calls (+6). - **GPT-OSS quality metrics look lower** (AdToolResult 0, Skill Invoked 0) but this matches baseline's already-low AdTool/Skill scores — not introduced by after-cohort hardening. **Cannot conclude** for Gemini 3.1 Pro or GPT-5.4 — failures are stack/connector infrastructure, not scored quality regressions. #### Harness notes - Stop competing Scout stacks before batch: `node scripts/evals stop` - `TEST_KIBANA_PORT`/`TEST_ES_PORT` patch in `base.config.ts` required when branch lacks `fix/weekly-evals-matrix` merge - Batch runner port-sync + full `boot_stack()` recovery landed in `run-security-evals-batch.sh` (skill-dev plugin, local) - Remaining gap: per-model connector-registration wait before `playwright.config.ts` loads (GPT-5.4 race) ### 6. Natural-routing convergence + OTLP trace evidence (`ad2-full-natural-20260713-191333`) **Harness change:** removed `configuration_overrides` / `AD2_EVAL_FORCE_SKILL_ROUTING` — default Agent Builder router only; `expectedSkills` is scoring-only. **Run:** `ad2-full-natural-20260713-191333` on `feat/search-skill-boundary` worktree (Claude 4.6 Sonnet, local Scout `:15001`/`:15000`). **9/9 Playwright passed** (25.3m). #### Representative trace cases (OTLP on local Scout ES — attach before stack restart) | Case | trace.id | Spans | Skill Invoked | ForbiddenTools | trajectory | Tool Calls | Quality | |------|----------|-------|---------------|----------------|------------|------------|---------| | **live-retrieval** | `bf1f541ffa9a8e91af088dd526a74ae6` | 13 | **1.0** | **1.0** | 0.75 | 4 | Criteria/Rubric/AdToolResult **1.0** | | **clean-profile (encoded-powershell)** | `f49eca1730170cf49fbf48b6a89fca0e` | 9 | **1.0** | **1.0** | 0.50 | 2 | Criteria/Rubric/AdToolResult **1.0** | **Live-retrieval OTLP tool sequence:** `load_skill` → `get_default_esql_query` (custom) → `platform.core.execute_esql` → `security.attack-discovery.run` (custom). No `generate_esql`, no `attack-discovery-alert-retrieval-builder`. **Clean-profile OTLP tool sequence:** `load_skill` → `security.attack-discovery.run` with provided alert IDs. **Skill-side fix (routing):** `attack_discovery_generator_skill.ts` — routing keywords, `ALERT_RETRIEVAL_BOUNDARIES` (minimal `get_default_esql_query` → `execute_esql` path), `STATUS_GUIDE` for status-only queries. **Local evidence artifacts** (operator worktree, gitignored): `.ad2-pr-evidence/ad2-full-natural-20260713-191333/` — `README.md`, `scores/score-summary.json`, `traces/*-otlp-projected.json`, raw OTLP JSON. ### Golden-cluster OTLP trace IDs (exported 2026-07-13) | Case | Golden `trace.id` | Run ID | Skill Invoked | ForbiddenTools | Waterfall | |------|-------------------|--------|---------------|----------------|-----------| | **live-retrieval** (happy path) | `7c29a6966e8766ffe343c9e3f8516aeb` | `ad2-golden-export-live-20260713-204938` | 1.0 | 1.0 | [APM trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=7c29a6966e8766ffe343c9e3f8516aeb) | | **clean-profile encoded-powershell** | `594f040b9f47df82f955107e6aaf5019` | `ad2-golden-export-20260713-203039` | 1.0 | 1.0 | [APM trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=594f040b9f47df82f955107e6aaf5019) | Local Scout OTLP snapshots (pre-teardown): `bf1f541ffa9a8e91af088dd526a74ae6`, `f49eca1730170cf49fbf48b6a89fca0e` — see `.ad2-pr-evidence/ad2-full-natural-20260713-191333/`. ### 7. Post-merge natural routing (2026-07-13) Latest commit on this branch drops harness forcing, adds clean-profile scenario-registry evals, and tunes trajectory precision to exclude `load_skill` under natural routing. ### 8. Skill-baseline golden-path scorecard (2026-07-15) **CI:** [Buildkite #469078](https://buildkite.com/elastic/kibana-pull-request/builds/469078) **PASS** on `208d3c253b0a` (skill prose reverted to upstream/main). **Local run:** `feat/attack-discovery-agent-builder-evals` @ `208d3c253b0a`, EIS Claude 4.6 Sonnet, `--grep "golden |non-golden"`, rep=1, Scout `evals_attack_discovery_agent_builder`. **5/5 Playwright passed** (11.7m). | Dataset | AdToolResult | Basic | Criteria | ForbiddenTools | Rubric | Skill Invoked | Tool Calls | WorkflowEvidence | Trajectory | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | provided-alerts | 1 | 1 | 1 | 1 | 1 | 1 | 4 | 1 | 0.33 | | live-retrieval (n=2) | 0 | 0.50 | **0** | **0** | 0.50 | 1 | **11.5** | — | **0.20** | | multiple-alert-sets | 0 | 1 | **0** | 1 | **0** | 1 | **9** | **0** | 0.14 | | missing-alert-retrieval | 0 | 1 | — | 0 | 1 | 1 | 9 | — | 0.29 | | status-only | 0 | 0 | — | 1 | 1 | 1 | 2 | — | 1 | **Interpretation:** With **main skill prose**, routing works (Skill Invoked = 1 everywhere) but **live-retrieval** and **multiple-alert-sets** show efficiency/quality gaps — extra tool calls, low Criteria/Rubric, missing WorkflowEvidence. These are the comparison targets for skill hardening in follow-up [patrykkopycinski#13](patrykkopycinski#13) (not in this PR). **Follow-up A/B (skill overlay from #13, same harness):** live-retrieval Criteria 0→1, Rubric 0.50→1, ForbiddenTools 0→1, tool calls 11.5→4, trajectory 0.20→0.83; multiple-alert-sets Criteria 0→1, Rubric 0→1, WorkflowEvidence 0→1, tool calls 9→2. See #13 for full Δ table. --------- Signed-off-by: Patryk Kopycinski <patryk.kopycinski@elastic.co> Co-authored-by: kibanamachine <42973632+kibanamachine@users.noreply.github.com>
qn895
pushed a commit
to qn895/kibana
that referenced
this pull request
Aug 11, 2026
…277625) ## Summary This PR introduces a dedicated, isolated evaluation suite for the **Attack Discovery 2.0 Agent Builder** integration. It is intentionally separate from the legacy direct-generation cohort so that we can measure workflow, routing, tool efficiency, and output quality independently. **The point of this PR is to show what focused evals buy you:** we did not just add tests — running them surfaced routing noise, wasted tool calls, and trace routing confusion, and those findings drove the **skill and harness** changes included here. **Baseline scope:** This PR ships the **eval suite + weekly CI wiring only**. `attack-discovery-generator` skill prose is intentionally left at **upstream/main** so follow-up PRs can measure improvement against this baseline. Committed evals use the **default Agent Builder router** — no `configuration_overrides`. Dataset `expectedSkills` is scoring-only (Skill Invoked evaluator). --- ## Eval profiles and CI cadence This suite ships **two automated cohorts** in one package; they answer different questions and run on different cadences. | Profile | Spec | Seed | CI cadence | Primary question | | --- | --- | --- | --- | --- | | **Golden-path** | `attack_discovery_agent_builder.spec.ts` | `fixtures.ts` (2 marker alerts) | **Weekly** — `llm_evals.yml` sets `EVAL_GREP: 'golden |non-golden'` | Default-agent **routing**, **tool/workflow plumbing**, efficiency gates | | **Clean profile** | `clean_profile_provided_alerts.spec.ts` | `scenario_registry/` (4 chains, 16 alerts + raw events) | **On-demand** — run with `--grep "clean profile"` or full suite locally/Buildkite on-demand | **Discovery quality** on realistic multi-stage chains (portable-seeder `clean` parity) | | **Full profile** | — | `ad-2.0-portable-seeder.py --profile full` | **Not automated** in this PR | Signal vs **noise** (~150+ distractor alerts); needs FPR/discrimination evaluators first | **Do not merge golden-path and clean profile for weekly CI** — golden-path includes live-retrieval, missing-retrieval, and status-only cases that depend on the minimal marker fixture; clean profile is provided-alerts-only on richer chain data. ## Evidence collected ### 1. Baseline (pre-prompt-tightening, first AD2 golden runs) | Path | Tool calls | Input tokens | Notes | |------|-----------|--------------|-------| | **provided-alerts** | 5 | ~266k | Skill invoked, AD tool ran | | **live-retrieval** | 13 | ~721k | Extra retrieval/corroboration tools | **Problem:** live-retrieval was doing 2.6× the tool calls and 2.7× the tokens for the same 2-alert fixture. Quality was fine, but efficiency was not. ### 2. After prompt tightening + new evaluators (clean stack restart) | Path | Tool calls | Input tokens | AdToolResult | Basic | Criteria | Rubric | Skill Invoked | Workflow | Trajectory | |------|-----------|--------------|--------------|-------|----------|--------|---------------|----------|------------| | **provided-alerts** | **4** | ~166k | 1 | 1 | 1 | 1 | 1 | 1 | 1 | | **live-retrieval** | **10** | ~387k | 1 | 1 | 1 | 1 | 1 | 1 | ~0.30 | **Improvement vs baseline:** | Metric | provided-alerts | live-retrieval | |--------|-----------------|----------------| | Tool calls | 5 → **4** (−20%) | 13 → **10** (−23%) | | Input tokens | ~266k → **~166k** (−38%) | ~721k → **~387k** (−46%) | | Quality evaluators | already 1.0 | already 1.0 | Trajectory on live-retrieval is intentionally low (~0.30) because the agent legitimately runs `get_default_esql_query` → `execute_esql` → `run`, but still adds corroboration detours. The new `ForbiddenTools` and `CostPerAlert` evaluators make those detours visible instead of hiding behind a quality score of 1. ### 3. Live-retrieval variance study (3 consecutive repetitions) Even after tightening, live-retrieval is **non-deterministic**: | Rep | Tool calls | Input tokens | Latency | ForbiddenTools | CostPerAlert | Quality | |-----|-----------|--------------|---------|----------------|--------------|---------| | 1 | 9 | 413k | 37s | 0 | 1 | 1 | | 2 | 14 | 657k | 135s | 1 | 0 | 1 | | 3 | 11 | 586k | 38s | 0 | 0 | 1 | | **mean** | **11.3** | **552k** | **70s** | **0.33** | **0.67** | **1.0** | - **ForbiddenTools** still leaks occasionally (`platform.core.get_document_by_id`, `platform.core.generate_esql`). - **CostPerAlert** fails 2/3 runs (threshold: 5 calls/alert × 2 alerts = 10; rep 2 hit 14). - **Quality stays at 1.0** across all reps — efficiency is the remaining gap, not output correctness. ### 4. What the evals revealed (and what we changed because of it) | Finding | Evidence | Change in this PR | |---------|----------|-------------------| | Default router sometimes picks wrong skill | Live-retrieval misroutes (baseline skill prose) | **Follow-up PR** — skill hardening moved out so this PR stays baseline | | Speculative tool calls inflate cost | 9–14 calls for 2 alerts; `generate_esql` / `get_document_by_id` detours | **Follow-up PR** — skill prompt tightening (baseline measures current behavior) | | Trace evaluators returned `n/a` | Agent Builder spans on local Scout ES, not golden cluster | Suite overrides `traceEsClient` to local Scout ES | | EIS connector setup flaky | Dot-prefixed `inferenceId` rejected; manual cleanup between runs | Documented as operator setup issue — **not** fixed in this PR (see out of scope) | | Quality-only scoring hides waste | Rubric = 1 while trajectory = 0.30 and CostPerAlert = 0 | Added `ForbiddenTools`, `CostPerAlert`, strict `Trajectory` evaluators | | Only happy paths tested | Two golden cases insufficient for regression signal | Added non-golden golden-path cases; clean-profile quality cohort kept **on-demand** (not weekly) | --- ## What changed ### New evals package `x-pack/solutions/security/packages/kbn-evals-suite-attack-discovery-agent-builder` - Deterministic fixtures with **relative timestamps** (static `@timestamp` caused live-retrieval to fail once alerts fell outside the "last six hours" window). - Custom evaluators: `WorkflowEvidence`, `StrictTrajectory`, `ResponseSkillInvocation`, `ForbiddenTools`, `CostPerAlert`, `AttackDiscoveryBasic`, `Criteria`, `Rubric`. - `chat_client.ts` uses natural routing only — no harness `configuration_overrides`; `expectedSkills` feeds Skill Invoked scoring. - Trajectory precision excludes `load_skill` (framework routing cost under natural routing). ### Dataset coverage | Path | Type | What it validates | |------|------|-------------------| | `provided-alerts` | Golden | Alerts attached; agent skips retrieval | | `live-retrieval` | Golden | Agent retrieves via default ES\|QL query | | `multiple-alert-sets` | Non-golden | Multiple alert attachments in one request | | `missing-alert-retrieval` | Non-golden | Live retrieval returns zero alerts; agent stops cleanly | | `status-only` | Non-golden | User asks for `execution_uuid` status; no new generation | | `clean-profile` (scenario registry) | On-demand | Four portable-seeder chains; **quality** evals (not weekly gate) | ### Infrastructure - Scout config set: `evals_attack_discovery_agent_builder` (AD 2.0 feature flags). - Suite registered in `evals.suites.json`, `package.json`, `tsconfig.base.json`. - Weekly pipeline: `llm_evals.yml` step for `attack-discovery-agent-builder` with `EVAL_GREP` (golden-path only). Clean profile remains in-repo for on-demand runs. ### Product improvements driven by evals - **None in this PR** — skill-side routing prose changes moved to the follow-up PR so this merge establishes a reproducible **main-baseline** scorecard. ### Intentionally out of scope This PR stays focused on the **AD2 eval suite**, **skill-side routing prose**, and **weekly/on-demand CI registration**. The following were explored during development but **reverted** to avoid unrelated platform diffs: - **`platform.core.generate_esql` tool description** — prefer-a-skill boundary clause is a separate Agent Builder platform change, not part of this evals PR. - **`evals_tracing` EIS `inferenceId` normalization** — connector-cache dot-prefix is an operator/local-setup concern; no Scout config change ships here. - **`@kbn/evals` `getToolCallSteps` `params` export** — the suite reads `load_skill` arguments via local `getToolCallStepsWithParams()` instead of extending the shared helper. --- ## How to run ```bash node scripts/evals run --suite attack-discovery-agent-builder --connector eis-anthropic-claude-5-sonnet ``` Focused development: ```bash node scripts/evals run --suite attack-discovery-agent-builder --grep "golden provided-alerts" --repetitions 1 ``` --- ## Follow-up (not blocking this PR) - Further reduce live-retrieval variance (currently 9–14 tool calls; CostPerAlert threshold is 5/alert). - Scout ES stack stability under consecutive heavy runs (infrastructure, not eval logic). - Rubric judge flake on golden encoded-powershell export (`judge_failed` / orphaned `tool_result` — quality evaluators still pass). ## Related - Skill-dev plugin lessons captured in [elastic/agent-builder-skill-dev-cursor-plugin#149](elastic/agent-builder-skill-dev-cursor-plugin#149) (`agent-builder-eval-playbook` skill). - Does **not** merge with the legacy direct-generation cohort (`kbn-evals-suite-attack-discovery`). ### 5. Weekly-matrix multi-model comparison (baseline vs after) Local batch runs on **2026-07-13** (M4, `feat/search-skill-boundary` worktree). LLM-gateway models skipped (`AD2_SKIP_GATEWAY_MODELS=true`). | Cohort | TEST_RUN_ID | Playwright result | |--------|-------------|-------------------| | baseline | `ad2-baseline-20260713-112322` | **6/6 models — 5/5 tests each** (`upstream/main` skill + `generate_esql`) | | after (initial) | `ad2-after-20260713-144634` | **3/6 complete** (Sonnet, Opus, Flash); 3 infra-failed | | after (retry) | `ad2-after-retry-20260713-153105` | **1/3 recovered** (GPT-OSS 5/5); Gemini 3.1 Pro + GPT-5.4 still infra-failed | #### Infrastructure fix applied before retry Dead-stack recovery was steering Playwright at `:5620` because `.scout/servers/local.json` was clobbered by competing Scout boots while the batch worker stayed on `:15001`/`:15000`. Batch runner now: 1. **`sync_scout_local_json()`** — patches (or recreates) `local.json` hosts before every suite dispatch and after every boot/reboot. 2. **`ensure_stack_alive()`** — uses full two-pass `boot_stack()` on dead-stack detection (not single-pass `boot_stack_once`), so EIS connectors are re-baked after ES wipe. #### Playwright pass matrix (merged after cohort) | Model | Baseline | After (merged) | Notes | |-------|----------|----------------|-------| | Claude 4.6 Sonnet | 5/5 | **5/5** | after initial run | | Claude 4.6 Opus | 5/5 | **5/5** | after initial run | | Gemini 3.0 Flash | 5/5 | **5/5** | after initial run | | Gemini 3.1 Pro | 5/5 | **2/5 fail** (retry also 2/5) | Stack died on `:15001` during converse; later tests hit missing `local.json` | | GPT-5.4 | 5/5 | **0/5** | Connector race at `playwright.config.ts` load (`Evaluation connector id … was not found`) — both initial + retry | | GPT-OSS-120b | 5/5 | **5/5** | Retry succeeded on `:15001`/`:15000` after port-sync fix | #### Evaluator Δ table (Overall means, log-parsed) Scores extracted from Playwright summary tables in batch logs (ephemeral Scout ES on `:15000` was torn down before ES-indexed export; `generate-ad2-comparison-report.mjs` could not query `.evaluation-scores*` post-teardown). **Δ = after − baseline.** | Model | PW | AdToolResult | WorkflowEvidence | trajectory | ForbiddenTools | AttackDiscoveryBasic | Criteria | Rubric | Skill Invoked | Tool Calls | Latency | |-------|-----|-------------:|-----------------:|-----------:|---------------:|---------------------:|---------:|-------:|--------------:|-----------:|--------:| | Claude 4.6 Sonnet | 5/5→5/5 | 0.50→0.50 (0) | 0.83→0.67 (−0.16) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 0.50→0.52 (+0.02) | 6.7→6.3 | 34.2s→31.4s | | Claude 4.6 Opus | 5/5→5/5 | 0.83→0.67 (−0.16) | 0.50→0.83 (+0.33) | 0.80→0.75 (−0.05) | 0.33→1 (+0.67) | 0.50→0.67 (+0.17) | 1→1 (0) | 0.33→0.83 (+0.50) | 0.24→0.37 (+0.13) | 18.3→12.0 | 60.5s→42.2s | | Gemini 3.0 Flash | 5/5→5/5 | 0.50→0.50 (0) | 0.50→0.50 (0) | 0.67→0.33 (−0.34) | 0.93→0.60 (−0.33) | 1→0.67 (−0.33) | 1→1 (0) | 1→0.67 (−0.33) | 0.16→0.23 (+0.07) | 15.2→21.2 | 13.2s→19.6s | | Gemini 3.1 Pro | 5/5→**2/5 fail** | — | — | — | — | — | — | — | — | — | — | | GPT-5.4 | 5/5→**0/5** | — | — | — | — | — | — | — | — | — | — | | GPT-OSS-120b | 5/5→5/5 | 0.50→0.00 (−0.50) | 0.17→0.17 (0) | 1→— | 0.33→0.40 (+0.07) | 1→0.83 (−0.17) | 1→0.83 (−0.17) | 0.67→0.83 (+0.16) | 0.43→0.00 (−0.43) | 5.2→12.2 | 52.9s→30.4s | Baseline GPT-OSS row from `ad2-baseline-20260713-112322` run log; after GPT-OSS row from retry run `ad2-after-retry-20260713-153105`. #### Regression watch (completed models only) On the **4 models with full 5/5 after Playwright completion** (Sonnet, Opus, Flash, GPT-OSS): - **No Criteria / AttackDiscoveryBasic regressions** — Criteria stays at 1.0 where scored; Basic dips on Flash (−0.33) and GPT-OSS (−0.17) but remains ≥0.67. - **Efficiency mixed:** Opus improves (fewer tool calls, lower latency); Flash regresses on trajectory (−0.34) and tool calls (+6). - **GPT-OSS quality metrics look lower** (AdToolResult 0, Skill Invoked 0) but this matches baseline's already-low AdTool/Skill scores — not introduced by after-cohort hardening. **Cannot conclude** for Gemini 3.1 Pro or GPT-5.4 — failures are stack/connector infrastructure, not scored quality regressions. #### Harness notes - Stop competing Scout stacks before batch: `node scripts/evals stop` - `TEST_KIBANA_PORT`/`TEST_ES_PORT` patch in `base.config.ts` required when branch lacks `fix/weekly-evals-matrix` merge - Batch runner port-sync + full `boot_stack()` recovery landed in `run-security-evals-batch.sh` (skill-dev plugin, local) - Remaining gap: per-model connector-registration wait before `playwright.config.ts` loads (GPT-5.4 race) ### 6. Natural-routing convergence + OTLP trace evidence (`ad2-full-natural-20260713-191333`) **Harness change:** removed `configuration_overrides` / `AD2_EVAL_FORCE_SKILL_ROUTING` — default Agent Builder router only; `expectedSkills` is scoring-only. **Run:** `ad2-full-natural-20260713-191333` on `feat/search-skill-boundary` worktree (Claude 4.6 Sonnet, local Scout `:15001`/`:15000`). **9/9 Playwright passed** (25.3m). #### Representative trace cases (OTLP on local Scout ES — attach before stack restart) | Case | trace.id | Spans | Skill Invoked | ForbiddenTools | trajectory | Tool Calls | Quality | |------|----------|-------|---------------|----------------|------------|------------|---------| | **live-retrieval** | `bf1f541ffa9a8e91af088dd526a74ae6` | 13 | **1.0** | **1.0** | 0.75 | 4 | Criteria/Rubric/AdToolResult **1.0** | | **clean-profile (encoded-powershell)** | `f49eca1730170cf49fbf48b6a89fca0e` | 9 | **1.0** | **1.0** | 0.50 | 2 | Criteria/Rubric/AdToolResult **1.0** | **Live-retrieval OTLP tool sequence:** `load_skill` → `get_default_esql_query` (custom) → `platform.core.execute_esql` → `security.attack-discovery.run` (custom). No `generate_esql`, no `attack-discovery-alert-retrieval-builder`. **Clean-profile OTLP tool sequence:** `load_skill` → `security.attack-discovery.run` with provided alert IDs. **Skill-side fix (routing):** `attack_discovery_generator_skill.ts` — routing keywords, `ALERT_RETRIEVAL_BOUNDARIES` (minimal `get_default_esql_query` → `execute_esql` path), `STATUS_GUIDE` for status-only queries. **Local evidence artifacts** (operator worktree, gitignored): `.ad2-pr-evidence/ad2-full-natural-20260713-191333/` — `README.md`, `scores/score-summary.json`, `traces/*-otlp-projected.json`, raw OTLP JSON. ### Golden-cluster OTLP trace IDs (exported 2026-07-13) | Case | Golden `trace.id` | Run ID | Skill Invoked | ForbiddenTools | Waterfall | |------|-------------------|--------|---------------|----------------|-----------| | **live-retrieval** (happy path) | `7c29a6966e8766ffe343c9e3f8516aeb` | `ad2-golden-export-live-20260713-204938` | 1.0 | 1.0 | [APM trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=7c29a6966e8766ffe343c9e3f8516aeb) | | **clean-profile encoded-powershell** | `594f040b9f47df82f955107e6aaf5019` | `ad2-golden-export-20260713-203039` | 1.0 | 1.0 | [APM trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=594f040b9f47df82f955107e6aaf5019) | Local Scout OTLP snapshots (pre-teardown): `bf1f541ffa9a8e91af088dd526a74ae6`, `f49eca1730170cf49fbf48b6a89fca0e` — see `.ad2-pr-evidence/ad2-full-natural-20260713-191333/`. ### 7. Post-merge natural routing (2026-07-13) Latest commit on this branch drops harness forcing, adds clean-profile scenario-registry evals, and tunes trajectory precision to exclude `load_skill` under natural routing. ### 8. Skill-baseline golden-path scorecard (2026-07-15) **CI:** [Buildkite #469078](https://buildkite.com/elastic/kibana-pull-request/builds/469078) **PASS** on `208d3c253b0a` (skill prose reverted to upstream/main). **Local run:** `feat/attack-discovery-agent-builder-evals` @ `208d3c253b0a`, EIS Claude 4.6 Sonnet, `--grep "golden |non-golden"`, rep=1, Scout `evals_attack_discovery_agent_builder`. **5/5 Playwright passed** (11.7m). | Dataset | AdToolResult | Basic | Criteria | ForbiddenTools | Rubric | Skill Invoked | Tool Calls | WorkflowEvidence | Trajectory | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | provided-alerts | 1 | 1 | 1 | 1 | 1 | 1 | 4 | 1 | 0.33 | | live-retrieval (n=2) | 0 | 0.50 | **0** | **0** | 0.50 | 1 | **11.5** | — | **0.20** | | multiple-alert-sets | 0 | 1 | **0** | 1 | **0** | 1 | **9** | **0** | 0.14 | | missing-alert-retrieval | 0 | 1 | — | 0 | 1 | 1 | 9 | — | 0.29 | | status-only | 0 | 0 | — | 1 | 1 | 1 | 2 | — | 1 | **Interpretation:** With **main skill prose**, routing works (Skill Invoked = 1 everywhere) but **live-retrieval** and **multiple-alert-sets** show efficiency/quality gaps — extra tool calls, low Criteria/Rubric, missing WorkflowEvidence. These are the comparison targets for skill hardening in follow-up [patrykkopycinski#13](patrykkopycinski#13) (not in this PR). **Follow-up A/B (skill overlay from elastic#13, same harness):** live-retrieval Criteria 0→1, Rubric 0.50→1, ForbiddenTools 0→1, tool calls 11.5→4, trajectory 0.20→0.83; multiple-alert-sets Criteria 0→1, Rubric 0→1, WorkflowEvidence 0→1, tool calls 9→2. See elastic#13 for full Δ table. --------- Signed-off-by: Patryk Kopycinski <patryk.kopycinski@elastic.co> Co-authored-by: kibanamachine <42973632+kibanamachine@users.noreply.github.com>
patrykkopycinski
added a commit
that referenced
this pull request
Aug 18, 2026
…277625) ## Summary This PR introduces a dedicated, isolated evaluation suite for the **Attack Discovery 2.0 Agent Builder** integration. It is intentionally separate from the legacy direct-generation cohort so that we can measure workflow, routing, tool efficiency, and output quality independently. **The point of this PR is to show what focused evals buy you:** we did not just add tests — running them surfaced routing noise, wasted tool calls, and trace routing confusion, and those findings drove the **skill and harness** changes included here. **Baseline scope:** This PR ships the **eval suite + weekly CI wiring only**. `attack-discovery-generator` skill prose is intentionally left at **upstream/main** so follow-up PRs can measure improvement against this baseline. Committed evals use the **default Agent Builder router** — no `configuration_overrides`. Dataset `expectedSkills` is scoring-only (Skill Invoked evaluator). --- ## Eval profiles and CI cadence This suite ships **two automated cohorts** in one package; they answer different questions and run on different cadences. | Profile | Spec | Seed | CI cadence | Primary question | | --- | --- | --- | --- | --- | | **Golden-path** | `attack_discovery_agent_builder.spec.ts` | `fixtures.ts` (2 marker alerts) | **Weekly** — `llm_evals.yml` sets `EVAL_GREP: 'golden |non-golden'` | Default-agent **routing**, **tool/workflow plumbing**, efficiency gates | | **Clean profile** | `clean_profile_provided_alerts.spec.ts` | `scenario_registry/` (4 chains, 16 alerts + raw events) | **On-demand** — run with `--grep "clean profile"` or full suite locally/Buildkite on-demand | **Discovery quality** on realistic multi-stage chains (portable-seeder `clean` parity) | | **Full profile** | — | `ad-2.0-portable-seeder.py --profile full` | **Not automated** in this PR | Signal vs **noise** (~150+ distractor alerts); needs FPR/discrimination evaluators first | **Do not merge golden-path and clean profile for weekly CI** — golden-path includes live-retrieval, missing-retrieval, and status-only cases that depend on the minimal marker fixture; clean profile is provided-alerts-only on richer chain data. ## Evidence collected ### 1. Baseline (pre-prompt-tightening, first AD2 golden runs) | Path | Tool calls | Input tokens | Notes | |------|-----------|--------------|-------| | **provided-alerts** | 5 | ~266k | Skill invoked, AD tool ran | | **live-retrieval** | 13 | ~721k | Extra retrieval/corroboration tools | **Problem:** live-retrieval was doing 2.6× the tool calls and 2.7× the tokens for the same 2-alert fixture. Quality was fine, but efficiency was not. ### 2. After prompt tightening + new evaluators (clean stack restart) | Path | Tool calls | Input tokens | AdToolResult | Basic | Criteria | Rubric | Skill Invoked | Workflow | Trajectory | |------|-----------|--------------|--------------|-------|----------|--------|---------------|----------|------------| | **provided-alerts** | **4** | ~166k | 1 | 1 | 1 | 1 | 1 | 1 | 1 | | **live-retrieval** | **10** | ~387k | 1 | 1 | 1 | 1 | 1 | 1 | ~0.30 | **Improvement vs baseline:** | Metric | provided-alerts | live-retrieval | |--------|-----------------|----------------| | Tool calls | 5 → **4** (−20%) | 13 → **10** (−23%) | | Input tokens | ~266k → **~166k** (−38%) | ~721k → **~387k** (−46%) | | Quality evaluators | already 1.0 | already 1.0 | Trajectory on live-retrieval is intentionally low (~0.30) because the agent legitimately runs `get_default_esql_query` → `execute_esql` → `run`, but still adds corroboration detours. The new `ForbiddenTools` and `CostPerAlert` evaluators make those detours visible instead of hiding behind a quality score of 1. ### 3. Live-retrieval variance study (3 consecutive repetitions) Even after tightening, live-retrieval is **non-deterministic**: | Rep | Tool calls | Input tokens | Latency | ForbiddenTools | CostPerAlert | Quality | |-----|-----------|--------------|---------|----------------|--------------|---------| | 1 | 9 | 413k | 37s | 0 | 1 | 1 | | 2 | 14 | 657k | 135s | 1 | 0 | 1 | | 3 | 11 | 586k | 38s | 0 | 0 | 1 | | **mean** | **11.3** | **552k** | **70s** | **0.33** | **0.67** | **1.0** | - **ForbiddenTools** still leaks occasionally (`platform.core.get_document_by_id`, `platform.core.generate_esql`). - **CostPerAlert** fails 2/3 runs (threshold: 5 calls/alert × 2 alerts = 10; rep 2 hit 14). - **Quality stays at 1.0** across all reps — efficiency is the remaining gap, not output correctness. ### 4. What the evals revealed (and what we changed because of it) | Finding | Evidence | Change in this PR | |---------|----------|-------------------| | Default router sometimes picks wrong skill | Live-retrieval misroutes (baseline skill prose) | **Follow-up PR** — skill hardening moved out so this PR stays baseline | | Speculative tool calls inflate cost | 9–14 calls for 2 alerts; `generate_esql` / `get_document_by_id` detours | **Follow-up PR** — skill prompt tightening (baseline measures current behavior) | | Trace evaluators returned `n/a` | Agent Builder spans on local Scout ES, not golden cluster | Suite overrides `traceEsClient` to local Scout ES | | EIS connector setup flaky | Dot-prefixed `inferenceId` rejected; manual cleanup between runs | Documented as operator setup issue — **not** fixed in this PR (see out of scope) | | Quality-only scoring hides waste | Rubric = 1 while trajectory = 0.30 and CostPerAlert = 0 | Added `ForbiddenTools`, `CostPerAlert`, strict `Trajectory` evaluators | | Only happy paths tested | Two golden cases insufficient for regression signal | Added non-golden golden-path cases; clean-profile quality cohort kept **on-demand** (not weekly) | --- ## What changed ### New evals package `x-pack/solutions/security/packages/kbn-evals-suite-attack-discovery-agent-builder` - Deterministic fixtures with **relative timestamps** (static `@timestamp` caused live-retrieval to fail once alerts fell outside the "last six hours" window). - Custom evaluators: `WorkflowEvidence`, `StrictTrajectory`, `ResponseSkillInvocation`, `ForbiddenTools`, `CostPerAlert`, `AttackDiscoveryBasic`, `Criteria`, `Rubric`. - `chat_client.ts` uses natural routing only — no harness `configuration_overrides`; `expectedSkills` feeds Skill Invoked scoring. - Trajectory precision excludes `load_skill` (framework routing cost under natural routing). ### Dataset coverage | Path | Type | What it validates | |------|------|-------------------| | `provided-alerts` | Golden | Alerts attached; agent skips retrieval | | `live-retrieval` | Golden | Agent retrieves via default ES\|QL query | | `multiple-alert-sets` | Non-golden | Multiple alert attachments in one request | | `missing-alert-retrieval` | Non-golden | Live retrieval returns zero alerts; agent stops cleanly | | `status-only` | Non-golden | User asks for `execution_uuid` status; no new generation | | `clean-profile` (scenario registry) | On-demand | Four portable-seeder chains; **quality** evals (not weekly gate) | ### Infrastructure - Scout config set: `evals_attack_discovery_agent_builder` (AD 2.0 feature flags). - Suite registered in `evals.suites.json`, `package.json`, `tsconfig.base.json`. - Weekly pipeline: `llm_evals.yml` step for `attack-discovery-agent-builder` with `EVAL_GREP` (golden-path only). Clean profile remains in-repo for on-demand runs. ### Product improvements driven by evals - **None in this PR** — skill-side routing prose changes moved to the follow-up PR so this merge establishes a reproducible **main-baseline** scorecard. ### Intentionally out of scope This PR stays focused on the **AD2 eval suite**, **skill-side routing prose**, and **weekly/on-demand CI registration**. The following were explored during development but **reverted** to avoid unrelated platform diffs: - **`platform.core.generate_esql` tool description** — prefer-a-skill boundary clause is a separate Agent Builder platform change, not part of this evals PR. - **`evals_tracing` EIS `inferenceId` normalization** — connector-cache dot-prefix is an operator/local-setup concern; no Scout config change ships here. - **`@kbn/evals` `getToolCallSteps` `params` export** — the suite reads `load_skill` arguments via local `getToolCallStepsWithParams()` instead of extending the shared helper. --- ## How to run ```bash node scripts/evals run --suite attack-discovery-agent-builder --connector eis-anthropic-claude-5-sonnet ``` Focused development: ```bash node scripts/evals run --suite attack-discovery-agent-builder --grep "golden provided-alerts" --repetitions 1 ``` --- ## Follow-up (not blocking this PR) - Further reduce live-retrieval variance (currently 9–14 tool calls; CostPerAlert threshold is 5/alert). - Scout ES stack stability under consecutive heavy runs (infrastructure, not eval logic). - Rubric judge flake on golden encoded-powershell export (`judge_failed` / orphaned `tool_result` — quality evaluators still pass). ## Related - Skill-dev plugin lessons captured in [elastic/agent-builder-skill-dev-cursor-plugin#149](elastic/agent-builder-skill-dev-cursor-plugin#149) (`agent-builder-eval-playbook` skill). - Does **not** merge with the legacy direct-generation cohort (`kbn-evals-suite-attack-discovery`). ### 5. Weekly-matrix multi-model comparison (baseline vs after) Local batch runs on **2026-07-13** (M4, `feat/search-skill-boundary` worktree). LLM-gateway models skipped (`AD2_SKIP_GATEWAY_MODELS=true`). | Cohort | TEST_RUN_ID | Playwright result | |--------|-------------|-------------------| | baseline | `ad2-baseline-20260713-112322` | **6/6 models — 5/5 tests each** (`upstream/main` skill + `generate_esql`) | | after (initial) | `ad2-after-20260713-144634` | **3/6 complete** (Sonnet, Opus, Flash); 3 infra-failed | | after (retry) | `ad2-after-retry-20260713-153105` | **1/3 recovered** (GPT-OSS 5/5); Gemini 3.1 Pro + GPT-5.4 still infra-failed | #### Infrastructure fix applied before retry Dead-stack recovery was steering Playwright at `:5620` because `.scout/servers/local.json` was clobbered by competing Scout boots while the batch worker stayed on `:15001`/`:15000`. Batch runner now: 1. **`sync_scout_local_json()`** — patches (or recreates) `local.json` hosts before every suite dispatch and after every boot/reboot. 2. **`ensure_stack_alive()`** — uses full two-pass `boot_stack()` on dead-stack detection (not single-pass `boot_stack_once`), so EIS connectors are re-baked after ES wipe. #### Playwright pass matrix (merged after cohort) | Model | Baseline | After (merged) | Notes | |-------|----------|----------------|-------| | Claude 4.6 Sonnet | 5/5 | **5/5** | after initial run | | Claude 4.6 Opus | 5/5 | **5/5** | after initial run | | Gemini 3.0 Flash | 5/5 | **5/5** | after initial run | | Gemini 3.1 Pro | 5/5 | **2/5 fail** (retry also 2/5) | Stack died on `:15001` during converse; later tests hit missing `local.json` | | GPT-5.4 | 5/5 | **0/5** | Connector race at `playwright.config.ts` load (`Evaluation connector id … was not found`) — both initial + retry | | GPT-OSS-120b | 5/5 | **5/5** | Retry succeeded on `:15001`/`:15000` after port-sync fix | #### Evaluator Δ table (Overall means, log-parsed) Scores extracted from Playwright summary tables in batch logs (ephemeral Scout ES on `:15000` was torn down before ES-indexed export; `generate-ad2-comparison-report.mjs` could not query `.evaluation-scores*` post-teardown). **Δ = after − baseline.** | Model | PW | AdToolResult | WorkflowEvidence | trajectory | ForbiddenTools | AttackDiscoveryBasic | Criteria | Rubric | Skill Invoked | Tool Calls | Latency | |-------|-----|-------------:|-----------------:|-----------:|---------------:|---------------------:|---------:|-------:|--------------:|-----------:|--------:| | Claude 4.6 Sonnet | 5/5→5/5 | 0.50→0.50 (0) | 0.83→0.67 (−0.16) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 0.50→0.52 (+0.02) | 6.7→6.3 | 34.2s→31.4s | | Claude 4.6 Opus | 5/5→5/5 | 0.83→0.67 (−0.16) | 0.50→0.83 (+0.33) | 0.80→0.75 (−0.05) | 0.33→1 (+0.67) | 0.50→0.67 (+0.17) | 1→1 (0) | 0.33→0.83 (+0.50) | 0.24→0.37 (+0.13) | 18.3→12.0 | 60.5s→42.2s | | Gemini 3.0 Flash | 5/5→5/5 | 0.50→0.50 (0) | 0.50→0.50 (0) | 0.67→0.33 (−0.34) | 0.93→0.60 (−0.33) | 1→0.67 (−0.33) | 1→1 (0) | 1→0.67 (−0.33) | 0.16→0.23 (+0.07) | 15.2→21.2 | 13.2s→19.6s | | Gemini 3.1 Pro | 5/5→**2/5 fail** | — | — | — | — | — | — | — | — | — | — | | GPT-5.4 | 5/5→**0/5** | — | — | — | — | — | — | — | — | — | — | | GPT-OSS-120b | 5/5→5/5 | 0.50→0.00 (−0.50) | 0.17→0.17 (0) | 1→— | 0.33→0.40 (+0.07) | 1→0.83 (−0.17) | 1→0.83 (−0.17) | 0.67→0.83 (+0.16) | 0.43→0.00 (−0.43) | 5.2→12.2 | 52.9s→30.4s | Baseline GPT-OSS row from `ad2-baseline-20260713-112322` run log; after GPT-OSS row from retry run `ad2-after-retry-20260713-153105`. #### Regression watch (completed models only) On the **4 models with full 5/5 after Playwright completion** (Sonnet, Opus, Flash, GPT-OSS): - **No Criteria / AttackDiscoveryBasic regressions** — Criteria stays at 1.0 where scored; Basic dips on Flash (−0.33) and GPT-OSS (−0.17) but remains ≥0.67. - **Efficiency mixed:** Opus improves (fewer tool calls, lower latency); Flash regresses on trajectory (−0.34) and tool calls (+6). - **GPT-OSS quality metrics look lower** (AdToolResult 0, Skill Invoked 0) but this matches baseline's already-low AdTool/Skill scores — not introduced by after-cohort hardening. **Cannot conclude** for Gemini 3.1 Pro or GPT-5.4 — failures are stack/connector infrastructure, not scored quality regressions. #### Harness notes - Stop competing Scout stacks before batch: `node scripts/evals stop` - `TEST_KIBANA_PORT`/`TEST_ES_PORT` patch in `base.config.ts` required when branch lacks `fix/weekly-evals-matrix` merge - Batch runner port-sync + full `boot_stack()` recovery landed in `run-security-evals-batch.sh` (skill-dev plugin, local) - Remaining gap: per-model connector-registration wait before `playwright.config.ts` loads (GPT-5.4 race) ### 6. Natural-routing convergence + OTLP trace evidence (`ad2-full-natural-20260713-191333`) **Harness change:** removed `configuration_overrides` / `AD2_EVAL_FORCE_SKILL_ROUTING` — default Agent Builder router only; `expectedSkills` is scoring-only. **Run:** `ad2-full-natural-20260713-191333` on `feat/search-skill-boundary` worktree (Claude 4.6 Sonnet, local Scout `:15001`/`:15000`). **9/9 Playwright passed** (25.3m). #### Representative trace cases (OTLP on local Scout ES — attach before stack restart) | Case | trace.id | Spans | Skill Invoked | ForbiddenTools | trajectory | Tool Calls | Quality | |------|----------|-------|---------------|----------------|------------|------------|---------| | **live-retrieval** | `bf1f541ffa9a8e91af088dd526a74ae6` | 13 | **1.0** | **1.0** | 0.75 | 4 | Criteria/Rubric/AdToolResult **1.0** | | **clean-profile (encoded-powershell)** | `f49eca1730170cf49fbf48b6a89fca0e` | 9 | **1.0** | **1.0** | 0.50 | 2 | Criteria/Rubric/AdToolResult **1.0** | **Live-retrieval OTLP tool sequence:** `load_skill` → `get_default_esql_query` (custom) → `platform.core.execute_esql` → `security.attack-discovery.run` (custom). No `generate_esql`, no `attack-discovery-alert-retrieval-builder`. **Clean-profile OTLP tool sequence:** `load_skill` → `security.attack-discovery.run` with provided alert IDs. **Skill-side fix (routing):** `attack_discovery_generator_skill.ts` — routing keywords, `ALERT_RETRIEVAL_BOUNDARIES` (minimal `get_default_esql_query` → `execute_esql` path), `STATUS_GUIDE` for status-only queries. **Local evidence artifacts** (operator worktree, gitignored): `.ad2-pr-evidence/ad2-full-natural-20260713-191333/` — `README.md`, `scores/score-summary.json`, `traces/*-otlp-projected.json`, raw OTLP JSON. ### Golden-cluster OTLP trace IDs (exported 2026-07-13) | Case | Golden `trace.id` | Run ID | Skill Invoked | ForbiddenTools | Waterfall | |------|-------------------|--------|---------------|----------------|-----------| | **live-retrieval** (happy path) | `7c29a6966e8766ffe343c9e3f8516aeb` | `ad2-golden-export-live-20260713-204938` | 1.0 | 1.0 | [APM trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=7c29a6966e8766ffe343c9e3f8516aeb) | | **clean-profile encoded-powershell** | `594f040b9f47df82f955107e6aaf5019` | `ad2-golden-export-20260713-203039` | 1.0 | 1.0 | [APM trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=594f040b9f47df82f955107e6aaf5019) | Local Scout OTLP snapshots (pre-teardown): `bf1f541ffa9a8e91af088dd526a74ae6`, `f49eca1730170cf49fbf48b6a89fca0e` — see `.ad2-pr-evidence/ad2-full-natural-20260713-191333/`. ### 7. Post-merge natural routing (2026-07-13) Latest commit on this branch drops harness forcing, adds clean-profile scenario-registry evals, and tunes trajectory precision to exclude `load_skill` under natural routing. ### 8. Skill-baseline golden-path scorecard (2026-07-15) **CI:** [Buildkite #469078](https://buildkite.com/elastic/kibana-pull-request/builds/469078) **PASS** on `208d3c253b0a` (skill prose reverted to upstream/main). **Local run:** `feat/attack-discovery-agent-builder-evals` @ `208d3c253b0a`, EIS Claude 4.6 Sonnet, `--grep "golden |non-golden"`, rep=1, Scout `evals_attack_discovery_agent_builder`. **5/5 Playwright passed** (11.7m). | Dataset | AdToolResult | Basic | Criteria | ForbiddenTools | Rubric | Skill Invoked | Tool Calls | WorkflowEvidence | Trajectory | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | provided-alerts | 1 | 1 | 1 | 1 | 1 | 1 | 4 | 1 | 0.33 | | live-retrieval (n=2) | 0 | 0.50 | **0** | **0** | 0.50 | 1 | **11.5** | — | **0.20** | | multiple-alert-sets | 0 | 1 | **0** | 1 | **0** | 1 | **9** | **0** | 0.14 | | missing-alert-retrieval | 0 | 1 | — | 0 | 1 | 1 | 9 | — | 0.29 | | status-only | 0 | 0 | — | 1 | 1 | 1 | 2 | — | 1 | **Interpretation:** With **main skill prose**, routing works (Skill Invoked = 1 everywhere) but **live-retrieval** and **multiple-alert-sets** show efficiency/quality gaps — extra tool calls, low Criteria/Rubric, missing WorkflowEvidence. These are the comparison targets for skill hardening in follow-up [#13](#13) (not in this PR). **Follow-up A/B (skill overlay from #13, same harness):** live-retrieval Criteria 0→1, Rubric 0.50→1, ForbiddenTools 0→1, tool calls 11.5→4, trajectory 0.20→0.83; multiple-alert-sets Criteria 0→1, Rubric 0→1, WorkflowEvidence 0→1, tool calls 9→2. See #13 for full Δ table. --------- Signed-off-by: Patryk Kopycinski <patryk.kopycinski@elastic.co> Co-authored-by: kibanamachine <42973632+kibanamachine@users.noreply.github.com>
patrykkopycinski
added a commit
that referenced
this pull request
Aug 18, 2026
…277625) ## Summary This PR introduces a dedicated, isolated evaluation suite for the **Attack Discovery 2.0 Agent Builder** integration. It is intentionally separate from the legacy direct-generation cohort so that we can measure workflow, routing, tool efficiency, and output quality independently. **The point of this PR is to show what focused evals buy you:** we did not just add tests — running them surfaced routing noise, wasted tool calls, and trace routing confusion, and those findings drove the **skill and harness** changes included here. **Baseline scope:** This PR ships the **eval suite + weekly CI wiring only**. `attack-discovery-generator` skill prose is intentionally left at **upstream/main** so follow-up PRs can measure improvement against this baseline. Committed evals use the **default Agent Builder router** — no `configuration_overrides`. Dataset `expectedSkills` is scoring-only (Skill Invoked evaluator). --- ## Eval profiles and CI cadence This suite ships **two automated cohorts** in one package; they answer different questions and run on different cadences. | Profile | Spec | Seed | CI cadence | Primary question | | --- | --- | --- | --- | --- | | **Golden-path** | `attack_discovery_agent_builder.spec.ts` | `fixtures.ts` (2 marker alerts) | **Weekly** — `llm_evals.yml` sets `EVAL_GREP: 'golden |non-golden'` | Default-agent **routing**, **tool/workflow plumbing**, efficiency gates | | **Clean profile** | `clean_profile_provided_alerts.spec.ts` | `scenario_registry/` (4 chains, 16 alerts + raw events) | **On-demand** — run with `--grep "clean profile"` or full suite locally/Buildkite on-demand | **Discovery quality** on realistic multi-stage chains (portable-seeder `clean` parity) | | **Full profile** | — | `ad-2.0-portable-seeder.py --profile full` | **Not automated** in this PR | Signal vs **noise** (~150+ distractor alerts); needs FPR/discrimination evaluators first | **Do not merge golden-path and clean profile for weekly CI** — golden-path includes live-retrieval, missing-retrieval, and status-only cases that depend on the minimal marker fixture; clean profile is provided-alerts-only on richer chain data. ## Evidence collected ### 1. Baseline (pre-prompt-tightening, first AD2 golden runs) | Path | Tool calls | Input tokens | Notes | |------|-----------|--------------|-------| | **provided-alerts** | 5 | ~266k | Skill invoked, AD tool ran | | **live-retrieval** | 13 | ~721k | Extra retrieval/corroboration tools | **Problem:** live-retrieval was doing 2.6× the tool calls and 2.7× the tokens for the same 2-alert fixture. Quality was fine, but efficiency was not. ### 2. After prompt tightening + new evaluators (clean stack restart) | Path | Tool calls | Input tokens | AdToolResult | Basic | Criteria | Rubric | Skill Invoked | Workflow | Trajectory | |------|-----------|--------------|--------------|-------|----------|--------|---------------|----------|------------| | **provided-alerts** | **4** | ~166k | 1 | 1 | 1 | 1 | 1 | 1 | 1 | | **live-retrieval** | **10** | ~387k | 1 | 1 | 1 | 1 | 1 | 1 | ~0.30 | **Improvement vs baseline:** | Metric | provided-alerts | live-retrieval | |--------|-----------------|----------------| | Tool calls | 5 → **4** (−20%) | 13 → **10** (−23%) | | Input tokens | ~266k → **~166k** (−38%) | ~721k → **~387k** (−46%) | | Quality evaluators | already 1.0 | already 1.0 | Trajectory on live-retrieval is intentionally low (~0.30) because the agent legitimately runs `get_default_esql_query` → `execute_esql` → `run`, but still adds corroboration detours. The new `ForbiddenTools` and `CostPerAlert` evaluators make those detours visible instead of hiding behind a quality score of 1. ### 3. Live-retrieval variance study (3 consecutive repetitions) Even after tightening, live-retrieval is **non-deterministic**: | Rep | Tool calls | Input tokens | Latency | ForbiddenTools | CostPerAlert | Quality | |-----|-----------|--------------|---------|----------------|--------------|---------| | 1 | 9 | 413k | 37s | 0 | 1 | 1 | | 2 | 14 | 657k | 135s | 1 | 0 | 1 | | 3 | 11 | 586k | 38s | 0 | 0 | 1 | | **mean** | **11.3** | **552k** | **70s** | **0.33** | **0.67** | **1.0** | - **ForbiddenTools** still leaks occasionally (`platform.core.get_document_by_id`, `platform.core.generate_esql`). - **CostPerAlert** fails 2/3 runs (threshold: 5 calls/alert × 2 alerts = 10; rep 2 hit 14). - **Quality stays at 1.0** across all reps — efficiency is the remaining gap, not output correctness. ### 4. What the evals revealed (and what we changed because of it) | Finding | Evidence | Change in this PR | |---------|----------|-------------------| | Default router sometimes picks wrong skill | Live-retrieval misroutes (baseline skill prose) | **Follow-up PR** — skill hardening moved out so this PR stays baseline | | Speculative tool calls inflate cost | 9–14 calls for 2 alerts; `generate_esql` / `get_document_by_id` detours | **Follow-up PR** — skill prompt tightening (baseline measures current behavior) | | Trace evaluators returned `n/a` | Agent Builder spans on local Scout ES, not golden cluster | Suite overrides `traceEsClient` to local Scout ES | | EIS connector setup flaky | Dot-prefixed `inferenceId` rejected; manual cleanup between runs | Documented as operator setup issue — **not** fixed in this PR (see out of scope) | | Quality-only scoring hides waste | Rubric = 1 while trajectory = 0.30 and CostPerAlert = 0 | Added `ForbiddenTools`, `CostPerAlert`, strict `Trajectory` evaluators | | Only happy paths tested | Two golden cases insufficient for regression signal | Added non-golden golden-path cases; clean-profile quality cohort kept **on-demand** (not weekly) | --- ## What changed ### New evals package `x-pack/solutions/security/packages/kbn-evals-suite-attack-discovery-agent-builder` - Deterministic fixtures with **relative timestamps** (static `@timestamp` caused live-retrieval to fail once alerts fell outside the "last six hours" window). - Custom evaluators: `WorkflowEvidence`, `StrictTrajectory`, `ResponseSkillInvocation`, `ForbiddenTools`, `CostPerAlert`, `AttackDiscoveryBasic`, `Criteria`, `Rubric`. - `chat_client.ts` uses natural routing only — no harness `configuration_overrides`; `expectedSkills` feeds Skill Invoked scoring. - Trajectory precision excludes `load_skill` (framework routing cost under natural routing). ### Dataset coverage | Path | Type | What it validates | |------|------|-------------------| | `provided-alerts` | Golden | Alerts attached; agent skips retrieval | | `live-retrieval` | Golden | Agent retrieves via default ES\|QL query | | `multiple-alert-sets` | Non-golden | Multiple alert attachments in one request | | `missing-alert-retrieval` | Non-golden | Live retrieval returns zero alerts; agent stops cleanly | | `status-only` | Non-golden | User asks for `execution_uuid` status; no new generation | | `clean-profile` (scenario registry) | On-demand | Four portable-seeder chains; **quality** evals (not weekly gate) | ### Infrastructure - Scout config set: `evals_attack_discovery_agent_builder` (AD 2.0 feature flags). - Suite registered in `evals.suites.json`, `package.json`, `tsconfig.base.json`. - Weekly pipeline: `llm_evals.yml` step for `attack-discovery-agent-builder` with `EVAL_GREP` (golden-path only). Clean profile remains in-repo for on-demand runs. ### Product improvements driven by evals - **None in this PR** — skill-side routing prose changes moved to the follow-up PR so this merge establishes a reproducible **main-baseline** scorecard. ### Intentionally out of scope This PR stays focused on the **AD2 eval suite**, **skill-side routing prose**, and **weekly/on-demand CI registration**. The following were explored during development but **reverted** to avoid unrelated platform diffs: - **`platform.core.generate_esql` tool description** — prefer-a-skill boundary clause is a separate Agent Builder platform change, not part of this evals PR. - **`evals_tracing` EIS `inferenceId` normalization** — connector-cache dot-prefix is an operator/local-setup concern; no Scout config change ships here. - **`@kbn/evals` `getToolCallSteps` `params` export** — the suite reads `load_skill` arguments via local `getToolCallStepsWithParams()` instead of extending the shared helper. --- ## How to run ```bash node scripts/evals run --suite attack-discovery-agent-builder --connector eis-anthropic-claude-5-sonnet ``` Focused development: ```bash node scripts/evals run --suite attack-discovery-agent-builder --grep "golden provided-alerts" --repetitions 1 ``` --- ## Follow-up (not blocking this PR) - Further reduce live-retrieval variance (currently 9–14 tool calls; CostPerAlert threshold is 5/alert). - Scout ES stack stability under consecutive heavy runs (infrastructure, not eval logic). - Rubric judge flake on golden encoded-powershell export (`judge_failed` / orphaned `tool_result` — quality evaluators still pass). ## Related - Skill-dev plugin lessons captured in [elastic/agent-builder-skill-dev-cursor-plugin#149](elastic/agent-builder-skill-dev-cursor-plugin#149) (`agent-builder-eval-playbook` skill). - Does **not** merge with the legacy direct-generation cohort (`kbn-evals-suite-attack-discovery`). ### 5. Weekly-matrix multi-model comparison (baseline vs after) Local batch runs on **2026-07-13** (M4, `feat/search-skill-boundary` worktree). LLM-gateway models skipped (`AD2_SKIP_GATEWAY_MODELS=true`). | Cohort | TEST_RUN_ID | Playwright result | |--------|-------------|-------------------| | baseline | `ad2-baseline-20260713-112322` | **6/6 models — 5/5 tests each** (`upstream/main` skill + `generate_esql`) | | after (initial) | `ad2-after-20260713-144634` | **3/6 complete** (Sonnet, Opus, Flash); 3 infra-failed | | after (retry) | `ad2-after-retry-20260713-153105` | **1/3 recovered** (GPT-OSS 5/5); Gemini 3.1 Pro + GPT-5.4 still infra-failed | #### Infrastructure fix applied before retry Dead-stack recovery was steering Playwright at `:5620` because `.scout/servers/local.json` was clobbered by competing Scout boots while the batch worker stayed on `:15001`/`:15000`. Batch runner now: 1. **`sync_scout_local_json()`** — patches (or recreates) `local.json` hosts before every suite dispatch and after every boot/reboot. 2. **`ensure_stack_alive()`** — uses full two-pass `boot_stack()` on dead-stack detection (not single-pass `boot_stack_once`), so EIS connectors are re-baked after ES wipe. #### Playwright pass matrix (merged after cohort) | Model | Baseline | After (merged) | Notes | |-------|----------|----------------|-------| | Claude 4.6 Sonnet | 5/5 | **5/5** | after initial run | | Claude 4.6 Opus | 5/5 | **5/5** | after initial run | | Gemini 3.0 Flash | 5/5 | **5/5** | after initial run | | Gemini 3.1 Pro | 5/5 | **2/5 fail** (retry also 2/5) | Stack died on `:15001` during converse; later tests hit missing `local.json` | | GPT-5.4 | 5/5 | **0/5** | Connector race at `playwright.config.ts` load (`Evaluation connector id … was not found`) — both initial + retry | | GPT-OSS-120b | 5/5 | **5/5** | Retry succeeded on `:15001`/`:15000` after port-sync fix | #### Evaluator Δ table (Overall means, log-parsed) Scores extracted from Playwright summary tables in batch logs (ephemeral Scout ES on `:15000` was torn down before ES-indexed export; `generate-ad2-comparison-report.mjs` could not query `.evaluation-scores*` post-teardown). **Δ = after − baseline.** | Model | PW | AdToolResult | WorkflowEvidence | trajectory | ForbiddenTools | AttackDiscoveryBasic | Criteria | Rubric | Skill Invoked | Tool Calls | Latency | |-------|-----|-------------:|-----------------:|-----------:|---------------:|---------------------:|---------:|-------:|--------------:|-----------:|--------:| | Claude 4.6 Sonnet | 5/5→5/5 | 0.50→0.50 (0) | 0.83→0.67 (−0.16) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 1→1 (0) | 0.50→0.52 (+0.02) | 6.7→6.3 | 34.2s→31.4s | | Claude 4.6 Opus | 5/5→5/5 | 0.83→0.67 (−0.16) | 0.50→0.83 (+0.33) | 0.80→0.75 (−0.05) | 0.33→1 (+0.67) | 0.50→0.67 (+0.17) | 1→1 (0) | 0.33→0.83 (+0.50) | 0.24→0.37 (+0.13) | 18.3→12.0 | 60.5s→42.2s | | Gemini 3.0 Flash | 5/5→5/5 | 0.50→0.50 (0) | 0.50→0.50 (0) | 0.67→0.33 (−0.34) | 0.93→0.60 (−0.33) | 1→0.67 (−0.33) | 1→1 (0) | 1→0.67 (−0.33) | 0.16→0.23 (+0.07) | 15.2→21.2 | 13.2s→19.6s | | Gemini 3.1 Pro | 5/5→**2/5 fail** | — | — | — | — | — | — | — | — | — | — | | GPT-5.4 | 5/5→**0/5** | — | — | — | — | — | — | — | — | — | — | | GPT-OSS-120b | 5/5→5/5 | 0.50→0.00 (−0.50) | 0.17→0.17 (0) | 1→— | 0.33→0.40 (+0.07) | 1→0.83 (−0.17) | 1→0.83 (−0.17) | 0.67→0.83 (+0.16) | 0.43→0.00 (−0.43) | 5.2→12.2 | 52.9s→30.4s | Baseline GPT-OSS row from `ad2-baseline-20260713-112322` run log; after GPT-OSS row from retry run `ad2-after-retry-20260713-153105`. #### Regression watch (completed models only) On the **4 models with full 5/5 after Playwright completion** (Sonnet, Opus, Flash, GPT-OSS): - **No Criteria / AttackDiscoveryBasic regressions** — Criteria stays at 1.0 where scored; Basic dips on Flash (−0.33) and GPT-OSS (−0.17) but remains ≥0.67. - **Efficiency mixed:** Opus improves (fewer tool calls, lower latency); Flash regresses on trajectory (−0.34) and tool calls (+6). - **GPT-OSS quality metrics look lower** (AdToolResult 0, Skill Invoked 0) but this matches baseline's already-low AdTool/Skill scores — not introduced by after-cohort hardening. **Cannot conclude** for Gemini 3.1 Pro or GPT-5.4 — failures are stack/connector infrastructure, not scored quality regressions. #### Harness notes - Stop competing Scout stacks before batch: `node scripts/evals stop` - `TEST_KIBANA_PORT`/`TEST_ES_PORT` patch in `base.config.ts` required when branch lacks `fix/weekly-evals-matrix` merge - Batch runner port-sync + full `boot_stack()` recovery landed in `run-security-evals-batch.sh` (skill-dev plugin, local) - Remaining gap: per-model connector-registration wait before `playwright.config.ts` loads (GPT-5.4 race) ### 6. Natural-routing convergence + OTLP trace evidence (`ad2-full-natural-20260713-191333`) **Harness change:** removed `configuration_overrides` / `AD2_EVAL_FORCE_SKILL_ROUTING` — default Agent Builder router only; `expectedSkills` is scoring-only. **Run:** `ad2-full-natural-20260713-191333` on `feat/search-skill-boundary` worktree (Claude 4.6 Sonnet, local Scout `:15001`/`:15000`). **9/9 Playwright passed** (25.3m). #### Representative trace cases (OTLP on local Scout ES — attach before stack restart) | Case | trace.id | Spans | Skill Invoked | ForbiddenTools | trajectory | Tool Calls | Quality | |------|----------|-------|---------------|----------------|------------|------------|---------| | **live-retrieval** | `bf1f541ffa9a8e91af088dd526a74ae6` | 13 | **1.0** | **1.0** | 0.75 | 4 | Criteria/Rubric/AdToolResult **1.0** | | **clean-profile (encoded-powershell)** | `f49eca1730170cf49fbf48b6a89fca0e` | 9 | **1.0** | **1.0** | 0.50 | 2 | Criteria/Rubric/AdToolResult **1.0** | **Live-retrieval OTLP tool sequence:** `load_skill` → `get_default_esql_query` (custom) → `platform.core.execute_esql` → `security.attack-discovery.run` (custom). No `generate_esql`, no `attack-discovery-alert-retrieval-builder`. **Clean-profile OTLP tool sequence:** `load_skill` → `security.attack-discovery.run` with provided alert IDs. **Skill-side fix (routing):** `attack_discovery_generator_skill.ts` — routing keywords, `ALERT_RETRIEVAL_BOUNDARIES` (minimal `get_default_esql_query` → `execute_esql` path), `STATUS_GUIDE` for status-only queries. **Local evidence artifacts** (operator worktree, gitignored): `.ad2-pr-evidence/ad2-full-natural-20260713-191333/` — `README.md`, `scores/score-summary.json`, `traces/*-otlp-projected.json`, raw OTLP JSON. ### Golden-cluster OTLP trace IDs (exported 2026-07-13) | Case | Golden `trace.id` | Run ID | Skill Invoked | ForbiddenTools | Waterfall | |------|-------------------|--------|---------------|----------------|-----------| | **live-retrieval** (happy path) | `7c29a6966e8766ffe343c9e3f8516aeb` | `ad2-golden-export-live-20260713-204938` | 1.0 | 1.0 | [APM trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=7c29a6966e8766ffe343c9e3f8516aeb) | | **clean-profile encoded-powershell** | `594f040b9f47df82f955107e6aaf5019` | `ad2-golden-export-20260713-203039` | 1.0 | 1.0 | [APM trace](https://kbn-evals-serverless-ed035a.kb.us-central1.gcp.elastic.cloud/app/observability/traces/explorer?traceId=594f040b9f47df82f955107e6aaf5019) | Local Scout OTLP snapshots (pre-teardown): `bf1f541ffa9a8e91af088dd526a74ae6`, `f49eca1730170cf49fbf48b6a89fca0e` — see `.ad2-pr-evidence/ad2-full-natural-20260713-191333/`. ### 7. Post-merge natural routing (2026-07-13) Latest commit on this branch drops harness forcing, adds clean-profile scenario-registry evals, and tunes trajectory precision to exclude `load_skill` under natural routing. ### 8. Skill-baseline golden-path scorecard (2026-07-15) **CI:** [Buildkite #469078](https://buildkite.com/elastic/kibana-pull-request/builds/469078) **PASS** on `208d3c253b0a` (skill prose reverted to upstream/main). **Local run:** `feat/attack-discovery-agent-builder-evals` @ `208d3c253b0a`, EIS Claude 4.6 Sonnet, `--grep "golden |non-golden"`, rep=1, Scout `evals_attack_discovery_agent_builder`. **5/5 Playwright passed** (11.7m). | Dataset | AdToolResult | Basic | Criteria | ForbiddenTools | Rubric | Skill Invoked | Tool Calls | WorkflowEvidence | Trajectory | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | provided-alerts | 1 | 1 | 1 | 1 | 1 | 1 | 4 | 1 | 0.33 | | live-retrieval (n=2) | 0 | 0.50 | **0** | **0** | 0.50 | 1 | **11.5** | — | **0.20** | | multiple-alert-sets | 0 | 1 | **0** | 1 | **0** | 1 | **9** | **0** | 0.14 | | missing-alert-retrieval | 0 | 1 | — | 0 | 1 | 1 | 9 | — | 0.29 | | status-only | 0 | 0 | — | 1 | 1 | 1 | 2 | — | 1 | **Interpretation:** With **main skill prose**, routing works (Skill Invoked = 1 everywhere) but **live-retrieval** and **multiple-alert-sets** show efficiency/quality gaps — extra tool calls, low Criteria/Rubric, missing WorkflowEvidence. These are the comparison targets for skill hardening in follow-up [#13](#13) (not in this PR). **Follow-up A/B (skill overlay from #13, same harness):** live-retrieval Criteria 0→1, Rubric 0.50→1, ForbiddenTools 0→1, tool calls 11.5→4, trajectory 0.20→0.83; multiple-alert-sets Criteria 0→1, Rubric 0→1, WorkflowEvidence 0→1, tool calls 9→2. See #13 for full Δ table. --------- Signed-off-by: Patryk Kopycinski <patryk.kopycinski@elastic.co> Co-authored-by: kibanamachine <42973632+kibanamachine@users.noreply.github.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Follow-up to elastic/kibana#277625 — adds the full profile portable-seeder cohort as an on-demand eval layer with explicit FPR/discrimination gates, and restores
attack-discovery-generatorskill hardening moved out of elastic#277625 so baseline Δ is measurable.Does not change weekly CI (golden-path
EVAL_GREPunchanged).Eval profiles in this change
attack_discovery_agent_builder.spec.tsEVAL_GREP)clean_profile_provided_alerts.spec.tsfull_profile_discrimination.spec.tsWhat changed
fullprofile: 7 signal chains (clean + AWS/Azure/macOS) + deterministic 110 background + 40 Defender loud-cluster alerts.NoiseFalsePositive,DiscoveryCountCap,MinValidatedDiscovery.full_profile_discrimination.spec.ts— live-retrieval with forbidden noise alert IDs and discovery count cap.attack_discovery_generator_skill.ts): routing keywords,ALERT_RETRIEVAL_BOUNDARIES(minimalget_default_esql_query→execute_esqlpath), forbiddengenerate_esql/ alert-retrieval-builder misroutes, status-onlyget_statuspath.Local A/B findings — skill hardening vs main baseline (2026-07-15)
Controlled comparison on the same harness (
evals_attack_discovery_agent_builder, EIS Claude 4.6 Sonnet,--grep "golden |non-golden", rep=1):208d3c253b0a, main skill)d288ae72335f, hardened skill)Evaluator Δ (golden-path means)
Takeaway: Skill hardening fixes the two baseline weak spots (elastic#277625 §8) — live-retrieval misroutes/extra tooling and multiple-alert-sets quality gates — without breaking Playwright pass rate.
missing-alert-retrievalBasic dropped 1→0 on one rep; worth a focused rerun before claiming regression.How to run
Golden-path (weekly gate):
node scripts/evals run --suite attack-discovery-agent-builder --grep "golden |non-golden"Full profile (on-demand):
node scripts/evals run --suite attack-discovery-agent-builder --grep "full profile"Test plan
full_profile_seed.test.ts,noise_fpr_evaluator.test.ts--grep "full profile"with EIS connectorEVAL_GREPstill excludes full-profile cases