Skip to content

[Alerting v2] Final API changes - #292548

Merged
cnasikas merged 34 commits into
elastic:mainfrom
cnasikas:alerting_v2_api_final_audit
Sep 25, 2026
Merged

cnasikas merged 34 commits into
elastic:mainfrom
cnasikas:alerting_v2_api_final_audit

Conversation

@cnasikas

@cnasikas cnasikas commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

Summary

A final pass on breaking changes for the alerting v2 APIs. This PR touches a lot of files, but the changes are very minimal and usually just renames of variables. I did not break them into smaller PRs due to the pressure of timing.

Changes

  • Rename sort param to sort_field in rule execution history API
  • Rename bulk-get/bulk-create rules response key to items
  • Document tag action as replace, not add
  • Rename total_events to total on action policy executions
  • Drop unused total from the match action policies response
  • Document response counts as estimates capped at 10,000
  • Rename the outcome filter to outcomes on both execution history endpoints
  • Unify the execution history outcome vocabulary
  • Align alerting v2 API timestamp naming and validation
  • Decouple the rule change-history wire schema from the UI row type
  • Name the rule execution duration fields for their units
  • Constrain client-chosen ids to a URL- and log-safe character set
  • Align bounds for concepts bounded differently across the v2 API
  • Fix status codes that were wrong, contradicted, or inconsistent
  • Move rule change history out from under the rule
  • Fix tests

Fixes: https://github.com/elastic/rna-program/issues/1099

Checklist

Check the PR satisfies following conditions.

Reviewers should verify this PR satisfies this list as well.

cnasikas and others added 3 commits September 21, 2026 19:53
Rename from/to -> start_time/end_time on the rule executions list
endpoint, and start_date -> start_time (plus new end_time) on the
action policy execution history endpoint, so both APIs use the same
naming for a time lower/upper bound instead of three different terms.

Verification per commit:
- node scripts/type_check --project x-pack/platform/packages/shared/response-ops/alerting-v2-schemas/tsconfig.json -> 0 errors
- node scripts/type_check --project x-pack/platform/plugins/shared/alerting_v2/tsconfig.json -> 0 errors
- node scripts/eslint --fix <changed files> -> 0 errors
- node scripts/jest <changed schema/public/server test files> -> 125 + 30 + 114 passed
The rule execution history endpoint used sort while the rules,
action policies, and rule template list endpoints all use sort_field
for the same concept. Renamed sort -> sort_field across the schema,
server routes/clients/queries, and public hooks/components for
consistency (BREAKING).

Verification:
- node scripts/type_check --project x-pack/platform/packages/shared/response-ops/alerting-v2-schemas/tsconfig.json -> 0 errors
- node scripts/type_check --project x-pack/platform/plugins/shared/alerting_v2/tsconfig.json -> 0 errors
- node scripts/eslint --fix <changed files> -> 0 errors
- node scripts/jest <changed test files> -> 59 + 28 + 82 tests passed

Co-authored-by: Cursor <cursoragent@cursor.com>
The rules list endpoint (and every other list endpoint) returns its
collection under items, while the bulk-get and bulk-create rules
responses used rules for the same concept. Renamed rules -> items in
bulkGetRulesResponseSchema and bulkCreateRulesResponseSchema (and all
server/OAS/scout consumers) for consistency (BREAKING). Request body
keys (e.g. bulkCreateRulesRequestSchema.rules) are unchanged; they are
a distinct request-only concept.

Verification:
- node scripts/type_check --project x-pack/platform/packages/shared/response-ops/alerting-v2-schemas/tsconfig.json -> 0 errors
- node scripts/type_check --project x-pack/platform/plugins/shared/alerting_v2/tsconfig.json -> 0 errors
- node scripts/eslint --fix <changed files> -> 0 errors
- node scripts/jest <changed test files> -> 225 + 184 + 13 tests passed

Co-authored-by: Cursor <cursoragent@cursor.com>
@cnasikas cnasikas self-assigned this Sep 22, 2026
@cnasikas cnasikas added release_note:skip Skip the PR/issue when compiling release notes backport:skip This PR does not require backporting Team:ResponseOps Platform ResponseOps team (formerly the Cases and Alerting teams) t// Feature:AlertingV2 labels Sep 22, 2026
cnasikas and others added 17 commits September 22, 2026 09:35
Unifies expiry (alert_action_schema.ts) and snooze_expiry
(alert_episode_schema.ts, episode_attachment_schema.ts, and derived
UI/ES|QL fields) to snoozed_until across request/response schemas,
mappers, ES|QL queries, and UI components/hooks.

The .alert-actions data stream mapping keeps the legacy expiry field
name; alert_actions_client.ts now maps snoozed_until to expiry on
write and downstream ES|QL queries reading directly from storage
still alias to expiry, so no data migration is required.

Co-authored-by: Cursor <cursoragent@cursor.com>
`stateTransitionOperatorSchema` ('AND'/'OR') and
`matchedActionPolicyCategorySchema` ('catch-all') were the only enum
values in the API that broke from the snake_case convention used
everywhere else. Renamed to 'and'/'or' and 'catch_all'.

`catch_all` is server-computed and never persisted, so the rename is a
straight sweep through the schema, `action_policy_client`, and all UI/
test consumers.

`state_transition.pending_operator`/`recovering_operator` is persisted
verbatim in the rule's saved object today, with no translation layer.
Migrating that value would need a saved-object model version bump; to
avoid that, `rules_client/utils.ts` now translates between the
lowercase API value and the legacy uppercase stored value at the
create/update/read boundaries, so the SO schema and existing rules are
untouched. Rule templates store the same `rule` shape without schema
validation, so `rule_templates_client/utils.ts` normalizes legacy
uppercase operators when parsing a stored template into the API
response, keeping older templates loadable.

Verification:
- node scripts/type_check --project x-pack/platform/packages/shared/response-ops/alerting-v2-schemas/tsconfig.json
- node scripts/type_check --project x-pack/platform/packages/shared/response-ops/alerting-v2-rule-form/tsconfig.json
- node scripts/type_check --project x-pack/platform/plugins/shared/alerting_v2/tsconfig.json
- node scripts/eslint --fix <changed files> (0 errors)
- node scripts/jest x-pack/platform/packages/shared/response-ops/alerting-v2-schemas (638 passed)
- node scripts/jest x-pack/platform/packages/shared/response-ops/alerting-v2-rule-form (1431 passed)
- node scripts/jest x-pack/platform/plugins/shared/alerting_v2 (4421 passed)

Co-authored-by: Cursor <cursoragent@cursor.com>
The "_tag" alert action schema described adding tags to an episode,
but the query layer takes the last tag actions value verbatim
(LAST(tags, @timestamp) WHERE action_type == "tag"), replacing the
episodes tags rather than appending to them. Update the descriptions
to say replace and document that an empty array clears the tags.

Co-authored-by: Cursor <cursoragent@cursor.com>
…executions

The action policy execution history list response named its count
total_events while the sibling rule execution history response named
the same concept total. Both back tabs of the same page, so the UI had
to spell it differently per tab. Rename the API field and the internal
ListExecutionHistoryResult.totalEvents to match the rule side.

Verified: type_check on alerting-v2-schemas and alerting_v2 (0 errors),
jest on the schema, server client, routes and UI suites (257 tests
passing), eslint on all changed files (0 errors).

Co-authored-by: Cursor <cursoragent@cursor.com>
…schema

The execution history test mocks hand-typed their fixtures, so several
had drifted to camelCase (perPage, searchMatches) for a snake_case
response. The fixtures were never read by the assertions, so the tests
passed while mocking a shape the API never returns.

Build every fixture through a builder returning the real
ListPolicyExecutionHistoryResponse / ListExecutionHistoryResult, type
the api mocks as MockedFunction of the client method, and drop the
as any on the route client result.

Verified the drift is now a compile error: reintroducing one camelCase
key per hardened fixture produced 6 TS2561 errors in 6 files ("Did you
mean to write per_page?"). After reverting, type_check on alerting_v2
is clean, 123 jest tests pass, eslint reports 0 errors.

Co-authored-by: Cursor <cursoragent@cursor.com>
…es response

POST /internal/alerting/v2/action_policies/_match returned a space-wide
total that no caller reads. The rule form only renders items, and the rule
details artifacts subsection derives its count from items.length, then uses
evaluated_count and is_truncated for the "23+" stat and its hint. Those two
stay; total was dead weight on an internal route.

The client keeps reading the find total to compute is_truncated, it just no
longer echoes it. is_truncated no longer names the removed field in its
description.

Verified: eslint 0 errors, type_check clean on alerting_v2,
alerting-v2-schemas and alerting-v2-rule-form, 898 jest tests pass across
the schemas package, the compose_discover form and the affected plugin
client, route and artifacts suites.

Co-authored-by: Cursor <cursoragent@cursor.com>
…,000

Every list response described its count as "the total number of X matching",
but almost none of them are exact. The saved-object reads behind GET /rules,
/action_policies and /rule_templates, the event-log reads behind the rule
execution and change history endpoints, and the search_matches counts all
leave track_total_hits unset, so Elasticsearch stops counting at 10,000 and
the reported number saturates there. The rule execution read also discards
the relation from hits.total, so a capped count is indistinguishable from an
exact one.

Describe all of them with one shared caveat from common.ts. The action
policy execution history total is exact today, but it carries the same
wording so no client depends on exactness and the read stays free to drop
track_total_hits later. The three history totals had no description at all
and now get one.

Verified: eslint 0 errors, type_check clean on alerting_v2,
alerting-v2-schemas and alerting-v2-rule-form, 718 jest tests pass across
the schemas package and the event log service suites.

Co-authored-by: Cursor <cursoragent@cursor.com>
…execution history endpoints

GET /execution_history/rules and /execution_history/action_policies both took
a repeated query param named 'outcome' while every sibling repeated param is
plural (rule_ids, episode_ids, tags). Both accept either a single value or an
array, so the singular name mislabelled the contract.

The internal layers were already plural below the client, so each route was
carrying a rename hop to get there: the rule route destructured
'outcome: outcomes' and the action policy client passed 'outcomes: outcome'.
Both hops collapse into shorthand now.

The per-event response field stays singular, as do the ES 'event.outcome'
term filter, the data-view field names and the single-select UI filter state
that feeds the array param.

Verified: eslint 0 errors, type_check clean on alerting_v2,
alerting-v2-schemas and alerting-v2-rule-form. The compiler caught 14 of the
call sites, including the scout API spec, since its URL helper is typed
against the request schema. Three component tests asserted the outgoing
payload through an untyped service mock and were caught by jest instead.
6488 tests pass: 636 schemas, 4421 alerting_v2, 1431 rule form.

Co-authored-by: Cursor <cursoragent@cursor.com>
The policy execution history API reported the stored event.action values
(dispatched, throttled, dispatch_failed) while the rule execution history
API reported ECS outcomes (success, failure), so the same concept had two
vocabularies across two endpoints.

Policy outcomes are now success, throttled and failure. The dispatch
detail stays where it already lived: failure_reason names the cause of a
failure, so encoding "dispatch" in the outcome added nothing.

The event log keeps writing the original event.action values; the new
outcome module translates in both directions at the API boundary, and the
event log service now names its filter 'actions' to match the vocabulary
it actually speaks. Both maps are exhaustive over their key type, so a new
outcome fails the type check rather than silently falling through.

User-visible labels are unchanged: the table and filter still read
Dispatched, Throttled and Failed.

Verification:
- node scripts/eslint --fix <changed files> -> 0 errors
- node scripts/type_check --project x-pack/platform/packages/shared/response-ops/alerting-v2-schemas/tsconfig.json -> 0 errors
- node scripts/type_check --project x-pack/platform/plugins/shared/alerting_v2/tsconfig.json -> 0 errors
- node scripts/jest x-pack/platform/packages/shared/response-ops/alerting-v2-schemas -> 638 passed
- node scripts/jest x-pack/platform/plugins/shared/alerting_v2 -> 4427 passed

Co-authored-by: Cursor <cursoragent@cursor.com>
Rename the change-history row timestamp to created_at and validate every
API timestamp as an ISO datetime.

- Rule change history: timestamp -> created_at. The value is the instant
  the history record was written, not when the rule changed: the only
  caller of logRuleChanges omits an explicit timestamp, so the document
  @timestamp is taken at build time and equals event.created. The rule's
  own change instant remains available as updated_at on the snapshot.
  The public adapter maps created_at onto the change-history-ui timestamp
  field, so the UI package type is untouched.
- Tighten z.string() to z.iso.datetime() for created_at, updated_at,
  started_at, ended_at and dispatched_at across the rule, action policy,
  rule execution and policy execution schemas.
- Document the alert-event timestamp as the source occurrence time. The
  field keeps its name because it is not user facing.

Verification:
- node scripts/eslint --fix (14 changed files): no eslint errors found
- node scripts/type_check --project alerting-v2-schemas: exited with 0
- node scripts/type_check --project alerting_v2: exited with 0
- node scripts/jest alerting-v2-schemas: 18 suites, 638 tests passed
- node scripts/jest alerting_v2: 381 suites, 4427 tests passed

Co-authored-by: Cursor <cursoragent@cursor.com>
Two fields on the change-history row were shaped by the change-history-ui
ChangeHistoryListItem interface rather than by the API.

- Replace the opaque metadata bag with an explicit version field. The wire
  type was a record of unknown, but the server only ever wrote
  { version: object.sequence }, so the API could never narrow it later. The
  public adapter rebuilds metadata.version for the UI package, which reads
  it for version-distance telemetry.
- Type action as an enum instead of a free string. The write path records a
  closed set of five lifecycle actions, so clients can now switch on it.
  audit_actions.ts is constrained to the same vocabulary, so the write path
  cannot record an action the read contract does not declare. Because the
  platform store types event.action as a string, the read path narrows and
  falls back to unknown, keeping the audit row rather than dropping it.

The dead document.metadata fallback goes away with the bag: alerting v2
never passes metadata or tags when logging.

Verification:
- node scripts/eslint --fix (10 changed files): no eslint errors found
- node scripts/type_check --project alerting-v2-schemas: exited with 0
- node scripts/type_check --project alerting_v2: exited with 0
- node scripts/jest alerting-v2-schemas: 18 suites, 639 tests passed
- node scripts/jest alerting_v2: 381 suites, 4429 tests passed

Co-authored-by: Cursor <cursoragent@cursor.com>
GET /execution_history/rules returned bare integers for two millisecond
values, so a client had no way to know the unit from the payload.

- timings.duration -> timings.duration_ms
- timings.scheduled_delay -> timings.scheduled_delay_ms
- sort_field value duration -> duration_ms, so callers sort by the name
  they read in the response

Both values come from nanosToMillis over Task Manager's event.duration
and kibana.task.schedule_delay. The descriptions now state the unit, and
scheduled_delay_ms documents that it goes negative when a run starts
ahead of schedule, which the schema tests already covered.

The internal RuleExecutionTimings domain type keeps its unsuffixed names;
the route maps at the boundary. Its doc comment previously argued that
grouping under timings made the unit implicit and suffixes unnecessary,
which no longer describes the HTTP shape, so it now just states the unit.

The UI hook maps its own sort vocabulary through an explicit table rather
than a ternary, so an absent sortField cannot fall through to the
non-default branch.

Out of scope: duration on alertEpisodeSchema and episodeAttachmentDataSchema
is an ES|QL view column (EVAL duration = DATE_DIFF("ms", ...) in
esql_views/alert_episodes.ts), not an HTTP API field, and renaming it means
changing the view.

Verification:
- node scripts/eslint --fix (13 changed files): no eslint errors found
- node scripts/type_check --project alerting-v2-schemas: exited with 0
- node scripts/type_check --project alerting_v2: exited with 0
- node scripts/jest alerting-v2-schemas: 18 suites, 639 tests passed
- node scripts/jest alerting_v2: 381 suites, 4429 tests passed

Co-authored-by: Cursor <cursoragent@cursor.com>
Clients name their own rules and action policies (PUT /rules/{id}), which
stays: it gives Terraform providers a natural idempotency key, so a create
that times out can be retried without orphaning a resource. The validation
was any string up to 150 characters, so spaces, path separators, query
delimiters, emoji and non-NFC Unicode were all accepted, and two ids could
look identical while differing byte-wise.

Add entityIdSchema in the schemas package: the existing trim/length bounds
plus ^[a-zA-Z0-9_-]+$, with the length kept as zod bounds rather than folded
into the pattern so the error says which rule failed. The server-generated
default is a UUID v4, which satisfies it.

Four inline copies of the identifier schema now derive from it: ruleIdSchema,
bulkByIdsSchema.ids, and the rule and action policy route path params. The
duplication is what let the constraint drift in the first place.

Ids also carry ENTITY_ID_NOTE, which states what the audit asked for: ids are
permanent, ids are not tomb-stoned, and re-using a deleted id means the new
resource inherits the execution history, change history and alert episodes
recorded under it, because all three key on the id.

Left alone: episode ids and event-source ids are not client-chosen for a
resource, rule template ids are a server-owned namespace, and the workflow
destination id belongs to the workflows plugin.

Verification:
- node scripts/eslint --fix (7 changed files): no eslint errors found
- node scripts/type_check --project alerting-v2-schemas: exited with 0
- node scripts/type_check --project alerting_v2: exited with 0
- node scripts/jest alerting-v2-schemas: 18 suites, 654 tests passed
- node scripts/jest alerting_v2: 381 suites, 4429 tests passed

Co-authored-by: Cursor <cursoragent@cursor.com>
Section 10 of the GA audit: the same concept carried a different bound
depending on where it appeared. Two of those gaps were functional bugs,
not just inconsistency.

Tags: a policy matcher tag allowed 256 characters while a rule tag
allows 128, so any matcher tag over 128 could never match anything.
Matcher tags now share the rule's length bound. The _match_for_rule
body, which carries one rule's tags, reuses the shared tags schema
instead of its own 100 x 256 pair.

KQL: matcher.expression allows MAX_KQL_LENGTH (4096) when a policy is
saved, but the field-suggestions endpoint capped the same expression at
2048, so a valid expression could be stored and then break autocomplete.
Both now use MAX_KQL_LENGTH.

group_hash: bounded at 256 in alert actions and 1024 in episodes, yet
every producer emits a SHA-256 hex digest. A shared groupHashSchema
states that exactly.

Also: artifact id aligned to ID_MAX_LENGTH, time_field aligned to
MAX_FIELD_NAME_LENGTH, and per_page capped at a single MAX_PER_PAGE of
100 (rules allowed 1000; nothing in Kibana requested more than 100).

The audit's remaining two identifier entries need no change: page
validation already shares queryIntSchema across every list endpoint, and
rule.id on _match_for_rule is gone since that body now carries only tags.

Verification:
- node scripts/eslint (changed files): 0 errors
- node scripts/type_check --project alerting-v2-schemas: 0 errors
- node scripts/type_check --project alerting_v2: 0 errors
- node scripts/type_check --project alerting, significant_events,
  alerting-v2-{browser-shared,common-queries,episodes-ui,rule-form,utils},
  yaml-rule-editor: 0 errors
- node scripts/jest alerting-v2-schemas: 664 passed
- node scripts/jest alerting_v2: 4429 passed
- node scripts/jest alerting-v2-{episodes-ui,rule-form,browser-shared,
  common-queries,utils}: 2581 passed

Co-authored-by: Cursor <cursoragent@cursor.com>
Section 8 of the GA audit, six rows.

POST /rules/{id}/_run now answers 202, not 204. `runSoon` reschedules the
task rather than running it, so "triggered successfully" overstated the
guarantee; the description now says the run is accepted and that concurrent
requests collapse into one.

POST /action_policies/{id}/_update_api_key now answers 200 with the policy,
matching its four siblings, and its OAS no longer claims a body it never
sent. The bulk path calls a new private `rotateApiKey` directly so it does
not pay for a read whose result the bulk response has nowhere to put, which
is how every other bulk operation in both clients already behaves. The 400
that a malformed {id} genuinely produces is now declared on all five
action-policy state routes instead of just one.

A superseded episode is no longer a 404. The episode was found; it is the
caller's assumption that failed, so `_activate` / `_deactivate` return 409
with the same ALERT_EPISODE_NOT_LATEST code. Only those two verbs can
produce it, so they declare it themselves rather than the shared factory
guessing from the action type.

Acting on an episode already in the requested state is also 409 rather than
400 — the request is well formed, the state conflicts. The bulk paths now
absorb 409 alongside 400/404 so one such item still cannot abort a batch.
The audit-only verbs stay permissive by design; the README now says so.

IMMUTABLE_FIELDS_CHANGED was only ever a stale README row: the throw site,
the scout spec and the OAS description all agree on 409. PUT and PATCH keep
their different statuses because the failures differ — PUT sends a valid
field with a conflicting value, PATCH sends a field the strict schema does
not define.

Route-level validation failures now carry the same treeified details.errors
the domain parsers attach. The shared zod helper flattens the ZodError into
a string before onRequestValidationError can see it, so alerting v2 wraps it
locally and keeps the issues; `_sourceSchema` is preserved so OAS generation
is unaffected. INVALID_RULE_DATA and INVALID_ACTION_POLICY_DATA are
unreachable over HTTP and are now documented as in-process guards.

Verification: node scripts/eslint --fix <changed files>            0 errors
  node scripts/type_check --project .../alerting_v2    0 errors
  node scripts/type_check --project .../alerting-v2-schemas  0 errors
  node scripts/jest .../alerting_v2                    382 suites, 4433 tests
  node scripts/jest .../alerting-v2-schemas            18 suites, 664 tests
Co-authored-by: Cursor <cursoragent@cursor.com>
Change-history rows are written for a rule's deletion and are meant to
outlive it, so nesting them under `/rules/{id}` promises a parent that may
no longer exist. Address them as their own collection filtered by rule,
matching the shape execution history already uses:

  GET /internal/alerting/v2/change_history/rules?rule_id=&page=&per_page=
  GET /internal/alerting/v2/change_history/rules/{change_id}?rule_id=

`rule_id` stays on the detail route only because `getHistory` is the sole
read API `@kbn/change-history` exposes and it always filters on an object
id; tracked for removal in elastic#292882.

Also swaps the change-history query schemas onto the shared `ruleIdSchema`
instead of redeclaring a local string schema.

Verification:
- node scripts/eslint --fix $(git status --short | awk '{print $NF}') -> 0 errors
- node scripts/type_check --project x-pack/platform/plugins/shared/alerting_v2/tsconfig.json -> 0 errors
- node scripts/type_check --project x-pack/platform/packages/shared/response-ops/alerting-v2-schemas/tsconfig.json -> 0 errors
- node scripts/type_check --project x-pack/platform/packages/shared/response-ops/alerting-v2-constants/tsconfig.json -> 0 errors
- node scripts/jest x-pack/platform/plugins/shared/alerting_v2 -> 4433 passed, 382 suites
- node scripts/jest x-pack/platform/packages/shared/response-ops/alerting-v2-schemas -> 666 passed, 18 suites

Co-authored-by: Cursor <cursoragent@cursor.com>
Conflict resolutions:
- run_rule_route: both sides moved 204 -> 202; kept the audit's fuller
  description covering run coalescing.
- update_action_policy_api_key_route: kept the audit's 200 + body, which
  matches the handler; main had only reworded the stale 204.
- rule_event_fields_schema.test: combined main's rename with the audit's
  MAX_KQL_LENGTH bound (the merged source already carries both).
- matched_policy_reason{,.test}: took main, which includes the same
  catch-all -> catch_all rename plus a tag badge redesign.
- action_policy_client.test, match_action_policies_route.test,
  use_matched_action_policies.test, linked_action_policies_step.test:
  kept the audit's removal of `total` from the match response, and
  widened a new main test's `get` mock now that rotation reads twice.

Verification:
- node scripts/kbn bootstrap (main changed package.json/pnpm-lock.yaml)
- node scripts/eslint --fix <resolved files> -> 0 errors
- node scripts/type_check --project x-pack/platform/plugins/shared/alerting_v2/tsconfig.json -> 0 errors
- node scripts/type_check --project .../alerting-v2-{schemas,constants,rule-form,episodes-ui}/tsconfig.json -> 0 errors
- node scripts/jest x-pack/platform/plugins/shared/alerting_v2 -> 4425 passed, 381 suites
- node scripts/jest .../alerting-v2-schemas -> 656 passed
- node scripts/jest .../alerting-v2-rule-form -> 1397 passed
- node scripts/jest .../alerting-v2-episodes-ui -> 918 passed

Co-authored-by: Cursor <cursoragent@cursor.com>
@cnasikas
cnasikas marked this pull request as ready for review September 23, 2026 07:45
@cnasikas
cnasikas requested a review from a team as a code owner September 23, 2026 07:45
@infra-vault-gh-plugin-prod

Copy link
Copy Markdown
Contributor

Pinging @elastic/response-ops (Team:ResponseOps)

@botelastic botelastic Bot added Feature:Embedding Embedding content via iFrame Team:obs-signals-traces Team:One Workflow Team label for One Workflow (Workflow automation) labels Sep 24, 2026
@cnasikas
cnasikas marked this pull request as draft September 24, 2026 12:59
cnasikas and others added 2 commits September 24, 2026 16:19
Reconciles the execution history API with elastic#292287, which
landed the same unification on main while this branch was in review.

- Time bounds: `from`/`to` (theirs). Kibana's published OAS has no
  precedent for `start_time`/`end_time`, and their PR already delivered
  the rule/policy unification this branch was after.
- Sort: `sort_field`/`sort_order` (ours) on both endpoints. Kibana-wide
  usage is 20/20 for the pair over bare `sort`, and v2's other three
  list endpoints already use `sort_field`.
- Outcome filter: `outcomes` (ours) on both endpoints, matching every
  other v2 array filter (`rule_ids`, `episode_ids`, `tags`).

Kept from main: count-only `per_page=0`, `.default()` paging,
`error.stack_trace`, `search_matches.is_truncated`.
Kept from this branch: `duration_ms`/`scheduled_delay_ms`, the outcome
vocabulary, ISO datetime validation, and the typed test fixtures.

OAS output left at main's version; the docs bundle is regenerated by CI.

Co-authored-by: Cursor <cursoragent@cursor.com>
These are written automatically by `node scripts/eslint_with_types` and
are meant to be deleted when it exits; three of them were swept into an
earlier commit on this branch.

Co-authored-by: Cursor <cursoragent@cursor.com>
@cnasikas
cnasikas force-pushed the alerting_v2_api_final_audit branch from 5bb9fbc to c23096d Compare September 24, 2026 14:58
@cnasikas
cnasikas removed request for a team September 24, 2026 14:59
cnasikas and others added 5 commits September 24, 2026 18:10
Three schema conflicts, all where main's actor work met this branch's
API audit:

- `action_policy_response_schema.ts`: took main's removal of `auth`.
  It was present at the merge base, so the follow-up PR that was going
  to drop it has landed.
- `created_by` / `updated_by` on action policies and rules: took main's
  `actorSchema`. `created_at` / `updated_at` keep this branch's
  `z.iso.datetime()` in place of the looser `z.string()`.
- `index.ts` and the `rule_data_schema.ts` imports: union of both
  export lists, so main's `actorSchema` / `Actor` join this branch's
  `entityIdSchema`, `ENTITY_ID_NOTE`, and `groupHashSchema`.

Verified: node scripts/eslint <resolved files> -> 0 errors
  node scripts/type_check --project x-pack/platform/plugins/shared/alerting_v2 -> 0 errors
  node scripts/type_check --project .../alerting-v2-schemas -> 0 errors
  node scripts/type_check --project .../alerting-v2-rule-form -> 0 errors
  node scripts/type_check --project .../alerting-v2-episodes-ui -> 0 errors
  node scripts/jest x-pack/platform/plugins/shared/alerting_v2 -> 384 suites, 4484 tests
  node scripts/jest --config .../alerting-v2-schemas -> 19 suites, 674 tests
  node scripts/jest --config .../alerting-v2-episodes-ui -> 142 suites, 1014 tests
  node scripts/jest --config .../alerting-v2-rule-form -> 1399 tests
Co-authored-by: Cursor <cursoragent@cursor.com>
One conflict, additive on both sides: main added a 403 license-forbidden
response to `enable_action_policy_route.ts` while this branch added the
400. Kept both, ascending, matching what the sibling action policy routes
landed on when they auto-merged.

Verified: node scripts/eslint --no-cache <resolved file> -> 0 errors
  node scripts/type_check --project x-pack/platform/plugins/shared/alerting_v2 -> 0 errors
  node scripts/jest x-pack/platform/plugins/shared/alerting_v2 -> 389 suites, 4561 tests
Co-authored-by: Cursor <cursoragent@cursor.com>
Three conflicts, all in generated Scout API manifests
(action_policies, alerts, rules). Their sha1s, test ids and line
numbers are derived, so the files were regenerated rather than
hand-merged. This also clears the stale `start_time` / `end_time`
titles the alerts manifest still carried from before the execution
history rename.

Verified: node scripts/scout update-test-config-manifests -> 4 manifests, all under alerting_v2
  node scripts/type_check --project x-pack/platform/plugins/shared/alerting_v2 -> 0 errors
  node scripts/jest x-pack/platform/plugins/shared/alerting_v2 -> 4557 tests passed
    (one suite lost its worker to SIGSEGV; 0 failed tests, passes in isolation)
Co-authored-by: Cursor <cursoragent@cursor.com>
@kibanamachine

Copy link
Copy Markdown
Contributor

PR run: bk-01a0d754-7d6e-4f41-81d1-da2468db9775::smoke-tests::anthropic-claude-4.5-haiku | Baseline (main): bk-01a0ce26-d149-4c77-804a-a59461a126b3::smoke-tests::anthropic-claude-4.5-haiku
Baseline: commit 60ca081, yesterday
Significance threshold: p < 0.05

Summary
No significant regressions detected (5 evaluator comparisons).

View full comparison in UI | Refresh baseline against latest main (click Unblock in the eval build)

No significant changes (5 rows)
Dataset Evaluator N Mean (PR) Mean (main) Diff p-value Sig Outcome
smoke tests: es-snapshot-loader SnapshotRestored 1 1.00 1.00 0.00 - n/a -
smoke tests: llm-judge Criteria 1 1.00 1.00 0.00 - n/a -
smoke tests: score ingestion and code evaluator ContainsKibana 1 1.00 1.00 0.00 - n/a -
smoke tests: trace-retrieval Input Tokens 1 12.00 12.00 0.00 - n/a -
smoke tests: trace-retrieval Output Tokens 1 4.00 4.00 0.00 - n/a -

@kibanamachine

Copy link
Copy Markdown
Contributor

💛 Build succeeded, but was flaky

Failed CI Steps

Metrics [docs]

Async chunks

Total size of all lazy-loaded chunks that will be downloaded as the user navigates the app

id before after diff
alertingVTwo 975.3KB 975.9KB +672.0B

Page load bundle

Size of the bundles that are downloaded on every page load. Target size is below 100kb

id before after diff
alertingVTwo 21.0KB 21.0KB +24.0B
shared-packages 4.8MB 4.8MB +562.0B
total +586.0B
Unknown metric groups

total optimizer output size

id before after diff
all 64.5MB 64.5MB +1.2KB

warm start memory

id before after diff
post forced gc heap baseline - 871540252 +871540252
post forced gc heap delta - -3542288 -3542288
post forced gc heap delta standard deviation - 2420721 +2420721
post forced gc heap target - 867997964 +867997964
tail heap delta - -24222860 -24222860
total +1714193789

workflow yaml validation

id before after diff
case_response.yaml (27 steps, 66 vars)/e2e/collectAllVariables 0 1 +1
case_response.yaml (27 steps, 66 vars)/e2e/performComputation 11 13 +2
case_response.yaml (27 steps, 66 vars)/e2e/total 16 19 +3
case_response.yaml (27 steps, 66 vars)/e2e/validateVariables 2 3 +1
infosec_demo.yaml (150 steps, 270 vars)/e2e/performComputation 46 55 +9
infosec_demo.yaml (150 steps, 270 vars)/e2e/total 119 150 +31
infosec_demo.yaml (150 steps, 270 vars)/e2e/validateIfConditions 4 6 +2
infosec_demo.yaml (150 steps, 270 vars)/e2e/validateLiquidTemplate 5 6 +1
infosec_demo.yaml (150 steps, 270 vars)/e2e/validateVariables 43 56 +13
infosec_demo.yaml (150 steps, 270 vars)/validateIfConditions 4 6 +2
infosec_demo.yaml (150 steps, 270 vars)/validateVariables (270 vars) 40 54 +14
total +79

Test Failures

  • [job] [logs] Jest Tests #7 / Bulk actions menu shows disable + export but not enable when the selected workflow is enabled
  • [job] [logs] FTR Configs #35 / Journey[login] Login
  • [job] [logs] Scout Lane #29 - stateful-classic / default / local-stateful-classic - Osquery live query details - accepts live query creation request (permission check)
  • [job] [logs] Rule Management - Security Solution Cypress Tests #6 / Rules table - privileges securitySolutionRulesV1.all should be able to adjust snooze settings should be able to adjust snooze settings
  • [job] [logs] Rule Management - Security Solution Cypress Tests #6 / Rules table - privileges securitySolutionRulesV1.read should not be able to adjust snooze settings should not be able to adjust snooze settings

History

cc @cnasikas

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport:skip This PR does not require backporting Feature:AlertingV2 release_note:skip Skip the PR/issue when compiling release notes Team:ResponseOps Platform ResponseOps team (formerly the Cases and Alerting teams) t// v9.6.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants