Skip to content

[Nightshift] Migrate Semantic Memory - #291288

Merged
normzhou merged 109 commits into
mainfrom
migration/deductive
Oct 1, 2026
Merged

normzhou merged 109 commits into
mainfrom
migration/deductive

Conversation

@normzhou

@normzhou normzhou commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Relates to elastic/nightshift-program#1549.

Closes elastic/nightshift-program#1614

Summary

Brings Deductive-style Semantic Memory

Storage and retrieval

  • Plugin-owned, hidden index nightshift-semantic-memory (context.semantic is the ELSER recall lane), reached only through a lazy internal client. No raw index route is exposed and workflow users get no index privileges.
  • Space-scoped like Cortex: IDs are ${space_id}:memory_<slug> and every query filters attributes.space_id. Only the Nightshift investigator reads or writes Memory; the agent id is not stored on memory documents.
  • Retrieval combines lexical and semantic RRF, read-time exponential decay, and Thompson sampling for browse mode. Hydration returns the top 15.
  • All Memory operations wait on one cached index/template readiness promise. A failed init fails Memory clearly without blocking plugin startup.

What the agent sees

  • Pages are written to /workspace/memories/<slug>.md and contain only the memory text: no front matter, ids, or scores.
  • /workspace/memories/.index.json is an additive catalog of { path, updated_at }. Nothing unrelated is deleted.
  • Only pages newly written this turn are announced ("Potentially relevant memories retrieved this turn:", with path and last-updated date). Cortex materializes its full wiki but announces nothing.
  • The model generates one title per memory and the slug is derived from it.

Round context

  • A final compose step returns a memory-only <system_update> as model_context and workflow_context["nightshift.semantic_memory.recall"] = { version: 1, data: { recalled_ids } }. Agent Builder persists both on the triggering round. The UI shows only the user's message, and later turns replay the model context verbatim until compaction.
  • Post-round optimize reads the recalled IDs from that round, so it does not depend on sandbox state or race with concurrent turns.
  • On HITL resume (round_execution_index > 0), materialization and decision-tree hydration are skipped. Other beforeAgent hooks still run.

Optimize

Three LLM calls run after each round, plus a compaction call when a written memory is over budget. Each gets the user task and the investigation in order: agent notes and each evidence tool call with a result excerpt, with ERROR on failed calls. Only the critique also gets the final answer.

  • Evidence only: all three calls drop calls (with their results) that are not evidence about the customer's environment: the agent's own process tools (progress reports, todos, list_files, skill loading, sleep, conversation metadata, user questions, sandbox file writes and edits) and any read of pre-seeded knowledge (paths under /workspace/cortex/ or /workspace/decision-trees/, /workspace/elastic.md, /workspace/connectors.md, by file_path/path or named in a bash command). Memory reads under /workspace/memories/ and unknown tools are kept, so the critique still sees which memories were opened.
  • Tool results: the post-round hook's tool_results input (added by [Nightshift] Give Cortex optimize the investigator's tool results #294383) is passed to both Cortex and Memory. Memory attaches each result to its call by tool_call_id. It first tries to read the persisted round, which also has the agent's notes; that read still fails from the workflow's request (Agent Builder reports the conversation as not found), so in practice the transcript is built from the hook's calls and results. Without results, the critique falls back to tool-call parameters, and extraction is skipped unless at least one evidence call returned output: a memory must be backed by a tool result, not by the final answer alone.
  • Critique labels recalled memories useful or harmful (ranked, one evidence budget of 128,000 estimated tokens shared by the recalled pages). Unused means no label, opening a memory is not using it, and harmful means its content was wrong or misleading.
  • Extraction is a classifier: it proposes up to three topics. Each entry has only a title, keywords (stored as tags) and replaces (recalled memory ids it supersedes: none for a new memory, one for an update, several for a merge). It writes no content. Recalled memories are shown with id, title, last-updated time and content. Its scope: organizational context, component behavior and environment topology that a tool result shows; never generic knowledge, how the run investigated, absence, secrets, container internals, or facts only relevant to this request.
  • Writer (nightshift_memory_write) writes the content of every entry, new ones included, from the topic and keywords, the replaced memories (if any, integrating their facts that still hold), the evidence transcript, the run time and the other entries' titles (so facts stay on their own topic). It has the same scope rules, adds no advice, dates facts that can change, writes a stopped problem in the past tense, and returns empty content, which writes nothing, when the run establishes nothing durable about the topic. It has a 4,000-character budget. Content over it goes to a compaction call (nightshift_memory_compact) with the writer's instructions, its input, the draft and a word target set below the budget; content still over the budget, or whose compaction call fails, is truncated and stored (one attempt, no retry).
  • Recall key: a written memory's context is this round's user task. Recall matches the next round's task against it, so a similar future task lands close to it.
  • Which memories an entry replaces: the ones it names in replaces, plus a live memory already stored at the entry's id even if it wasn't recalled (extraction only sees recalled memories). An archived memory at that id is written over. There is no lexical or catalog overlap matching. A memory already replaced by an earlier entry in the round skips the later entry.
  • Harmful memories are archived up front with reason harmful and never reach the writer, even when an entry replaces them. That entry is written from only the live memories it replaces, or as a new memory.
  • Writes: one exact-slug create-or-merge flow. An existing live slug is merged in place, a missing slug is created, and a create 409 fetches the winner and merges. Source pages are archived (reason merged) only after the canonical write succeeds; the merge revalidates its sources around synthesis and resynthesizes on change. An id that belongs to an archived memory is reused: the new memory is written over the archived document under optimistic concurrency, with counters reset and merge history kept.
  • Impression/conversion counters use same-index optimistic concurrency (versioned read, JS decay, partial update guarded by _seq_no/_primary_term, bounded retries). This avoids scripts, which Elasticsearch forbids on indices with semantic_text.
  • Memory optimize resolves its model through the Nightshift resolver from [Nightshift] Remove model selection from Model Settings #293894 (@kbn/nightshift-ai), given the step's connector_id and the triggering round's round_connector_id, like Cortex. Both share one telemetry tag (significant_events_investigation), so spend is attributed to Nightshift. elastic/nightshift-program#1644 renames the tag.

Safety

  • Non-canonical memory IDs are rejected before building sandbox paths or notifications. Slug canonicalization is idempotent (every leading memory- prefix stripped, cut to 80 characters with no trailing hyphen), so every written id is canonical.
  • Secret detection covers the proposed title, keywords, the task, and the written content.
  • Materialize and optimize both skip unless the agent is the Nightshift investigator, while storage stays Space-scoped.
  • Memory steps stop with the workflow: cancellation or the step timeout aborts every Memory LLM call, and no later call or write starts.
  • Partial sandbox writes count as materialization failures. Optimize still runs when the sandbox is unavailable.
  • Persistent workflow logs never contain page IDs or slugs, connector IDs, prompts, or raw model errors.
  • Memory persistence failures are reported as failed optimizations.

Telemetry

  • Aggregate EBT events nightshift-semantic-memory-materialized and nightshift-semantic-memory-optimized carry correlation IDs, bounded outcomes, and counts only. Model, duration, and token details come from existing Agent Builder and workflow telemetry.

Managed workflows

  • system-nightshift-sandbox-materialize-workspace: obtain sandbox, then Cortex and Memory in parallel (single attempt, 45s fail-fast), then compose model and workflow context.
  • system-nightshift-agent-optimize: obtain sandbox, then Cortex and Memory optimize in settled parallel branches. It replaces the standalone Cortex optimize workflow, declares tool_calls and tool_results (Agent Builder only sends tool_results to workflows that declare it), and passes both to Cortex and Memory.
  • system-nightshift-decision-tree-hydrate: restores stored decision trees once per round for the investigation and reinforcement agents.

Known limitations

  • The writer integrates the replaced memories, so a line copied from Cortex before the evidence filter can persist through later updates. The filter stops new seeded content; it doesn't clean old memories.
  • An active topic is rewritten most rounds (refreshed counts and "last confirmed" dates), and a merge can archive a memory a later round would have recalled. Tracked for evals.
  • The writer sometimes keeps an unverified inference from a replaced memory or a stale present-tense sentence. These persist with tool results, so they are prompt-following issues, tracked for evals.
  • The seeded-path filter is a heuristic: it matches the known roots and docs by path parameter or bash command, not every way a tool could read them.

Test plan

  • CI: pending on cde9c6ad51d (after merging main and regenerating the workflow step schemas).
  • Nightshift plugin: 894 tests, run in-band on cde9c6ad51d. Covers the evidence-only filter (process tools, seeded paths by parameter and bash command, memory reads kept, critique and extraction inputs), extraction skipped without evidence output (empty, reasoning-only, empty-result and seeded-only rounds), the writer budget and compaction, slug idempotency and the truncation boundary, Space isolation, counters, budgets, merge and create-race recovery, OCC retries, source preservation, round-context handoff, telemetry, managed workflow wiring, the Memory transcript renderer (ordering, results, error marking, budgets, fallback, building from hook results), topic entries (new, update, merge, a live memory at the entry's id joining unnamed), the writer writing every entry and skipping empty content, the task as recall key, archived-id reuse, and harmful-memory handling.
  • Managed workflow definitions: 1,723 tests on cde9c6ad51d. This includes real-YAML checks that resume-only steps run at index 0 and not at 1 or 2, that both strict post-execution schemas declare the inputs Agent Builder sends (now including tool_results), and that both optimize branches receive tool_results as an array.
  • Real Elasticsearch integration (2 tests): the production mapping and hidden setting, lexical and browse retrieval, Space isolation, concurrent counter increments, and archiving.
  • Static checks: Nightshift type-check and lint pass; workflow step schemas regenerated.
  • Agent Builder: the generic round-context persistence, replay, HITL, and compaction behavior is covered in [Agent Builder] Pass model and workflow context through workflow hooks #292787 (merged) and not duplicated here.
  • LLM task audit (local, gitignored, reads the Phoenix trace of the optimizer's LLM calls, plus a written AI judgment of every call) on 5cf304cf8f1: the critique and extraction got every evidence call's result (6 of 19 calls were evidence) and none of the dropped calls; no Cortex, decision-tree or environment-doc content reached any Memory call; extraction and the writer got no final answer. The judgments found the limitations above.
  • Live end-to-end with a real Claude Sonnet connector, driving the Agent Builder API and managed workflows the UI uses (local, gitignored probes). It validates wiring and persistence, not answer quality.
    • On 3692d5ea962: round 1 26/26; the critique and extraction got only the 7 evidence calls of 20, the writer's prompt carried the budget and it wrote 2,822 characters (no compaction needed).
    • On 5cf304cf8f1: round 1 26/26; round 2 40/44, where the 4 failures are seed-specific checks after round 1 correctly merged the seed into a new memory. On f94bcc3629d (after the [Nightshift] Remove model selection from Model Settings #293894 merge): release checks 16/16 (optimize gets the round connector, spend tagged) and HITL 7/7.
    • On d548aa2f53c: two rounds, 26/26 each, both optimize branches completed. One round updated two memories in place; the other wrote a new memory through the writer. Every written memory's context was that round's task.
    • On a3e6f2fd159: two rounds, 27/27 and 25/25 assertions. The second round updated a memory in place with a recurrence found only in a query result.
    • On f74c1e369a6 (earlier head): two ordinary turns, 44/44 assertions (memory retrieved and materialized, announced on the first turn only, recalled IDs persisted per round, user message unchanged, hidden context replayed verbatim, no memory IDs in workflow logs, counters persisted); every optimize LLM call carried feature_id=significant_events_investigation (9/9); HITL resume 7/7 (materialize and tree hydrate ran only at index 0, optimize and reinforce each ran once with the round's recalled IDs and connector).

@cla-checker-service

cla-checker-service Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

💚 CLA has been signed

@kibanamachine kibanamachine added the reviewer:libra PR review with Libra. This disables Claude and Scout reviewers label Sep 15, 2026
@kibanamachine

Copy link
Copy Markdown
Contributor

Selected for Libra review

This PR was selected for Libra review as part of the temporary 50% trial.

To opt out permanently, remove the reviewer:libra label. It will not be added again to this PR.

@github-actions

github-actions Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

Elastic Docs Style Checker (Vale)

Summary: 5 suggestions found

💡 Suggestions (5): Optional style improvements. Apply when helpful.
File Line Rule Message
docs/reference/connectors-kibana/_snippets/data-context-sources-connectors-list.md 36 Elastic.WordChoice Consider using 'run, start' instead of 'execute', unless the term is in the UI.
docs/reference/connectors-kibana/mysql-action-type.md 4 Elastic.WordChoice Consider using 'run, start' instead of 'execute', unless the term is in the UI.
docs/reference/connectors-kibana/mysql-action-type.md 12 Elastic.WordChoice Consider using 'run, start' instead of 'Execute', unless the term is in the UI.
docs/reference/connectors-kibana/mysql-action-type.md 91 Elastic.WordChoice Consider using 'run, start' instead of 'Execute', unless the term is in the UI.
docs/reference/connectors-kibana/mysql-action-type.md 114 Elastic.WordChoice Consider using 'can, might' instead of 'may', unless the term is in the UI.

The Vale linter checks documentation changes against the Elastic Docs style guide. To use Vale locally or report issues, refer to Elastic style guide for Vale.

@normzhou
normzhou force-pushed the migration/deductive branch 2 times, most recently from e91d698 to d282958 Compare September 15, 2026 23:41
@normzhou
normzhou changed the base branch from main to nightshift-cortex-hydrate-optimize September 16, 2026 02:05
@normzhou
normzhou force-pushed the migration/deductive branch 2 times, most recently from f0fb598 to f5c213e Compare September 16, 2026 15:38
@normzhou
normzhou changed the base branch from nightshift-cortex-hydrate-optimize to main September 16, 2026 15:38
@normzhou
normzhou force-pushed the migration/deductive branch 2 times, most recently from c7d1f55 to 461c52b Compare September 19, 2026 06:21
@normzhou
normzhou changed the base branch from main to workflow-step-execution-logs September 19, 2026 06:21
@normzhou
normzhou changed the base branch from workflow-step-execution-logs to main September 22, 2026 15:26
@normzhou
normzhou changed the base branch from main to feature/agent-workflow-model-context September 22, 2026 18:12
@normzhou
normzhou force-pushed the feature/agent-workflow-model-context branch from bcafa89 to 7b2fd8a Compare September 22, 2026 18:28
@normzhou
normzhou marked this pull request as ready for review September 22, 2026 20:55
@normzhou
normzhou requested review from a team as code owners September 22, 2026 20:55
@kibanamachine

Copy link
Copy Markdown
Contributor

PR size reminder

This PR has 7047 added lines of reviewable code, which is above the 500-line guideline for Nightshift PRs.

Large PRs get significantly less review engagement and take longer to merge. Consider splitting this into smaller, focused PRs before requesting review.

Comment thread x-pack/solutions/observability/plugins/nightshift_investigations/server/plugin.ts Outdated
@botelastic botelastic Bot added the Team:One Workflow Team label for One Workflow (Workflow automation) label Sep 22, 2026
normzhou and others added 3 commits September 30, 2026 17:13
The combined optimize workflow now passes tool_results to the Memory step as
well as Cortex. Memory attaches them to each tool call and, when the persisted
round cannot be read, builds its investigation transcript from those calls and
results instead of falling back to parameters only.

Co-authored-by: Cursor <cursoragent@cursor.com>
…iter

Extraction now returns only each entry's title, keywords, and the recalled
memories it replaces. The writer writes the content of every entry, new ones
included, from the transcript and the replaced memories. The recall key is
this round's task. Lexical overlap matching is gone: an entry merges the
memories it names, plus a live memory already at its id.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
kibanamachine and others added 5 commits October 1, 2026 02:31
…writing

The writer and extractor now see only the task and the investigation; the final answer
can carry the agent's inferences. Extraction is skipped when no tool results are available.

Co-authored-by: Cursor <cursoragent@cursor.com>
…nd writing

Reads of the files Nightshift seeded into the sandbox, and the agent's progress reports that
restate them, are dropped from the transcript the extractor and writer see. The critique still
gets every call, since it judges whether recalled memories were used.

Co-authored-by: Cursor <cursoragent@cursor.com>
… too

The critique judges whether recalled memories were opened and followed, which the memory reads
show; the seeded reads and progress reports only lengthened its input.

Co-authored-by: Cursor <cursoragent@cursor.com>
Brings in #293894, which moved Nightshift model choice off Model Settings.

- Cortex and Memory optimize now pick their model with the Nightshift
  resolver inside createOptimizeModel: strict connector_id, then the
  round's model, then the Nightshift default. Model errors fail the
  branch instead of skipping it.
- agent_optimize and both optimize steps declare a bounded connector_id
  next to round_connector_id.
- decision_tree_reinforce keeps the resolve_model step from main, plus
  workflow_context from this branch.
- cortex_optimize workflow stays deleted (replaced by agent_optimize).
- Generated workflow step schemas regenerated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@achyutjhunjhunwala

Copy link
Copy Markdown
Contributor

@normzhou I merged main into this branch (f94bcc3) now that #293894 has landed. What changed on your side:

  • createOptimizeModel now picks the model with the Nightshift resolver first: strict connector_id, then round_connector_id, then the Nightshift default. It then loads that model through Agent Builder as before, with your telemetry helper. Model errors now throw instead of skipping, so a bad or blocked model fails the Cortex or Memory branch visibly. The only quiet skip left is a missing plugin.
  • agent_optimize and both optimize steps take a bounded connector_id next to round_connector_id.
  • Decision tree reinforce keeps the resolve_model step from main, plus your workflow_context input.
  • Tests for register_cortex, register_memory, create_optimize_model and both steps are updated. The generated step schemas were regenerated.

Pull before your next push. Happy to walk through any of it.

…ranscripts

Memory calls now also skip todos, file listings, skills, file edits, and reads of elastic.md and
connectors.md, and match seeded paths on any tool's path parameter. Unknown tools are kept.

Co-authored-by: Cursor <cursoragent@cursor.com>
Cutting a slug to 80 characters right after a hyphen produced an id that isCanonicalMemoryId
rejects, so the memory was stored but never recalled or updated.

Co-authored-by: Cursor <cursoragent@cursor.com>
Stripping the memory- prefix once meant a title like 'Memory Memory Memory Pressure' was
reduced again at each stage and stored under an id isCanonicalMemoryId rejects. Strip every
leading memory- prefix and hyphen in one pass.

Co-authored-by: Cursor <cursoragent@cursor.com>
normzhou and others added 4 commits October 1, 2026 11:58
An empty or reasoning-only round, empty result arrays, or results only from seeded reads and
process tools gave extraction no observed output, yet it could still propose memories.

Co-authored-by: Cursor <cursoragent@cursor.com>
The writer had no length budget, so long content was cut at 6,000 characters and the replaced
memories were still archived, losing facts that were only in the cut tail. The writer now has a
4,000-character budget. A memory over it goes to one compaction call with the writer's goal, its
input, the overage and a word target set below the budget. If compaction fails or is still too
long, the memory is truncated and stored; cancellation still stops the write.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	src/platform/packages/private/kbn-workflow-step-schema-cli/generated/index.json
#	src/platform/packages/private/kbn-workflow-step-schema-cli/generated/strict/schema.json
#	src/platform/packages/private/kbn-workflow-step-schema-cli/generated/template/schema.json
Co-authored-by: Cursor <cursoragent@cursor.com>
@normzhou
normzhou enabled auto-merge October 1, 2026 20:01
@kibanamachine

Copy link
Copy Markdown
Contributor

💛 Build succeeded, but was flaky

Failed CI Steps

Metrics [docs]

Unknown metric groups

warm start memory

id before after diff
post forced gc heap baseline - 889075618 +889075618
post forced gc heap delta - -1278684 -1278684
post forced gc heap delta standard deviation - 1883575 +1883575
post forced gc heap target - 887796934 +887796934
tail heap delta - -20934632 -20934632
total +1756542811

workflow yaml validation

id before after diff
case_response.yaml (27 steps, 66 vars)/e2e/performComputation 12 11 -1
case_response.yaml (27 steps, 66 vars)/e2e/total 17 16 -1
infosec_demo.yaml (150 steps, 270 vars)/e2e/performComputation 50 47 -3
infosec_demo.yaml (150 steps, 270 vars)/e2e/total 127 117 -10
infosec_demo.yaml (150 steps, 270 vars)/e2e/validateIfConditions 5 4 -1
infosec_demo.yaml (150 steps, 270 vars)/e2e/validateVariables 47 41 -6
infosec_demo.yaml (150 steps, 270 vars)/validateIfConditions 5 4 -1
infosec_demo.yaml (150 steps, 270 vars)/validateVariables (270 vars) 46 39 -7
total -30

Test Failures

  • [job] [logs] Jest Tests #6 / EvaluatorEditorFlyout testing before saving abandons the run when the draft changes during the profile probe
  • [job] [logs] Scout Lane #102 - serverless-security_complete / default / local-serverless-security_complete - Osquery live query details - accepts live query creation request (permission check)

History

@normzhou
normzhou added this pull request to the merge queue Oct 1, 2026
Merged via the queue into main with commit e4fba58 Oct 1, 2026
39 checks passed
@normzhou

normzhou commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

Post-merge bookkeeping: low-pri follow-up for the Known limitations / prompt-quality gaps from closed-loop testing — elastic/nightshift-program#1723 (nice-to-have). Not blocking; revisit only if Memory quality becomes a focus.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport:skip This PR does not require backporting release_note:skip Skip the PR/issue when compiling release notes reviewer:libra PR review with Libra. This disables Claude and Scout reviewers Team:One Workflow Team label for One Workflow (Workflow automation) v9.6.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants