Repository navigation
[Nightshift] Migrate Semantic Memory - #291288
Conversation
|
💚 CLA has been signed |
Selected for Libra reviewThis PR was selected for Libra review as part of the temporary 50% trial. To opt out permanently, remove the |
7781da2 to
cb9dcb6
Compare
Elastic Docs Style Checker (Vale)Summary: 5 suggestions found 💡 Suggestions (5): Optional style improvements. Apply when helpful.
The Vale linter checks documentation changes against the Elastic Docs style guide. To use Vale locally or report issues, refer to Elastic style guide for Vale. |
e91d698 to
d282958
Compare
f0fb598 to
f5c213e
Compare
c7d1f55 to
461c52b
Compare
461c52b to
cec34f2
Compare
cec34f2 to
121cf89
Compare
bcafa89 to
7b2fd8a
Compare
121cf89 to
72f9c0d
Compare
PR size reminderThis PR has 7047 added lines of reviewable code, which is above the 500-line guideline for Nightshift PRs. Large PRs get significantly less review engagement and take longer to merge. Consider splitting this into smaller, focused PRs before requesting review. |
The combined optimize workflow now passes tool_results to the Memory step as well as Cortex. Memory attaches them to each tool call and, when the persisted round cannot be read, builds its investigation transcript from those calls and results instead of falling back to parameters only. Co-authored-by: Cursor <cursoragent@cursor.com>
…iter Extraction now returns only each entry's title, keywords, and the recalled memories it replaces. The writer writes the content of every entry, new ones included, from the transcript and the replaced memories. The recall key is this round's task. Lexical overlap matching is gone: an entry merges the memories it names, plus a live memory already at its id. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…writing The writer and extractor now see only the task and the investigation; the final answer can carry the agent's inferences. Extraction is skipped when no tool results are available. Co-authored-by: Cursor <cursoragent@cursor.com>
…nd writing Reads of the files Nightshift seeded into the sandbox, and the agent's progress reports that restate them, are dropped from the transcript the extractor and writer see. The critique still gets every call, since it judges whether recalled memories were used. Co-authored-by: Cursor <cursoragent@cursor.com>
… too The critique judges whether recalled memories were opened and followed, which the memory reads show; the seeded reads and progress reports only lengthened its input. Co-authored-by: Cursor <cursoragent@cursor.com>
Brings in #293894, which moved Nightshift model choice off Model Settings. - Cortex and Memory optimize now pick their model with the Nightshift resolver inside createOptimizeModel: strict connector_id, then the round's model, then the Nightshift default. Model errors fail the branch instead of skipping it. - agent_optimize and both optimize steps declare a bounded connector_id next to round_connector_id. - decision_tree_reinforce keeps the resolve_model step from main, plus workflow_context from this branch. - cortex_optimize workflow stays deleted (replaced by agent_optimize). - Generated workflow step schemas regenerated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
@normzhou I merged
Pull before your next push. Happy to walk through any of it. |
…ranscripts Memory calls now also skip todos, file listings, skills, file edits, and reads of elastic.md and connectors.md, and match seeded paths on any tool's path parameter. Unknown tools are kept. Co-authored-by: Cursor <cursoragent@cursor.com>
Cutting a slug to 80 characters right after a hyphen produced an id that isCanonicalMemoryId rejects, so the memory was stored but never recalled or updated. Co-authored-by: Cursor <cursoragent@cursor.com>
Stripping the memory- prefix once meant a title like 'Memory Memory Memory Pressure' was reduced again at each stage and stored under an id isCanonicalMemoryId rejects. Strip every leading memory- prefix and hyphen in one pass. Co-authored-by: Cursor <cursoragent@cursor.com>
An empty or reasoning-only round, empty result arrays, or results only from seeded reads and process tools gave extraction no observed output, yet it could still propose memories. Co-authored-by: Cursor <cursoragent@cursor.com>
The writer had no length budget, so long content was cut at 6,000 characters and the replaced memories were still archived, losing facts that were only in the cut tail. The writer now has a 4,000-character budget. A memory over it goes to one compaction call with the writer's goal, its input, the overage and a word target set below the budget. If compaction fails or is still too long, the memory is truncated and stored; cancellation still stops the write. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # src/platform/packages/private/kbn-workflow-step-schema-cli/generated/index.json # src/platform/packages/private/kbn-workflow-step-schema-cli/generated/strict/schema.json # src/platform/packages/private/kbn-workflow-step-schema-cli/generated/template/schema.json
Co-authored-by: Cursor <cursoragent@cursor.com>
💛 Build succeeded, but was flaky
Failed CI StepsMetrics [docs]Unknown metric groupswarm start memory
workflow yaml validation
Test Failures
History
|
|
Post-merge bookkeeping: low-pri follow-up for the Known limitations / prompt-quality gaps from closed-loop testing — elastic/nightshift-program#1723 ( |
Relates to elastic/nightshift-program#1549.
Closes elastic/nightshift-program#1614
Summary
Brings Deductive-style Semantic Memory
Storage and retrieval
nightshift-semantic-memory(context.semanticis the ELSER recall lane), reached only through a lazy internal client. No raw index route is exposed and workflow users get no index privileges.${space_id}:memory_<slug>and every query filtersattributes.space_id. Only the Nightshift investigator reads or writes Memory; the agent id is not stored on memory documents.What the agent sees
/workspace/memories/<slug>.mdand contain only the memory text: no front matter, ids, or scores./workspace/memories/.index.jsonis an additive catalog of{ path, updated_at }. Nothing unrelated is deleted.Round context
<system_update>asmodel_contextandworkflow_context["nightshift.semantic_memory.recall"] = { version: 1, data: { recalled_ids } }. Agent Builder persists both on the triggering round. The UI shows only the user's message, and later turns replay the model context verbatim until compaction.round_execution_index> 0), materialization and decision-tree hydration are skipped. OtherbeforeAgenthooks still run.Optimize
Three LLM calls run after each round, plus a compaction call when a written memory is over budget. Each gets the user task and the investigation in order: agent notes and each evidence tool call with a result excerpt, with
ERRORon failed calls. Only the critique also gets the final answer.list_files, skill loading, sleep, conversation metadata, user questions, sandbox file writes and edits) and any read of pre-seeded knowledge (paths under/workspace/cortex/or/workspace/decision-trees/,/workspace/elastic.md,/workspace/connectors.md, byfile_path/pathor named in a bash command). Memory reads under/workspace/memories/and unknown tools are kept, so the critique still sees which memories were opened.tool_resultsinput (added by [Nightshift] Give Cortex optimize the investigator's tool results #294383) is passed to both Cortex and Memory. Memory attaches each result to its call bytool_call_id. It first tries to read the persisted round, which also has the agent's notes; that read still fails from the workflow's request (Agent Builder reports the conversation as not found), so in practice the transcript is built from the hook's calls and results. Without results, the critique falls back to tool-call parameters, and extraction is skipped unless at least one evidence call returned output: a memory must be backed by a tool result, not by the final answer alone.replaces(recalled memory ids it supersedes: none for a new memory, one for an update, several for a merge). It writes no content. Recalled memories are shown with id, title, last-updated time and content. Its scope: organizational context, component behavior and environment topology that a tool result shows; never generic knowledge, how the run investigated, absence, secrets, container internals, or facts only relevant to this request.nightshift_memory_write) writes the content of every entry, new ones included, from the topic and keywords, the replaced memories (if any, integrating their facts that still hold), the evidence transcript, the run time and the other entries' titles (so facts stay on their own topic). It has the same scope rules, adds no advice, dates facts that can change, writes a stopped problem in the past tense, and returns empty content, which writes nothing, when the run establishes nothing durable about the topic. It has a 4,000-character budget. Content over it goes to a compaction call (nightshift_memory_compact) with the writer's instructions, its input, the draft and a word target set below the budget; content still over the budget, or whose compaction call fails, is truncated and stored (one attempt, no retry).contextis this round's user task. Recall matches the next round's task against it, so a similar future task lands close to it.replaces, plus a live memory already stored at the entry's id even if it wasn't recalled (extraction only sees recalled memories). An archived memory at that id is written over. There is no lexical or catalog overlap matching. A memory already replaced by an earlier entry in the round skips the later entry.harmfuland never reach the writer, even when an entry replaces them. That entry is written from only the live memories it replaces, or as a new memory.409fetches the winner and merges. Source pages are archived (reasonmerged) only after the canonical write succeeds; the merge revalidates its sources around synthesis and resynthesizes on change. An id that belongs to an archived memory is reused: the new memory is written over the archived document under optimistic concurrency, with counters reset and merge history kept._seq_no/_primary_term, bounded retries). This avoids scripts, which Elasticsearch forbids on indices withsemantic_text.@kbn/nightshift-ai), given the step'sconnector_idand the triggering round'sround_connector_id, like Cortex. Both share one telemetry tag (significant_events_investigation), so spend is attributed to Nightshift. elastic/nightshift-program#1644 renames the tag.Safety
memory-prefix stripped, cut to 80 characters with no trailing hyphen), so every written id is canonical.Telemetry
nightshift-semantic-memory-materializedandnightshift-semantic-memory-optimizedcarry correlation IDs, bounded outcomes, and counts only. Model, duration, and token details come from existing Agent Builder and workflow telemetry.Managed workflows
system-nightshift-sandbox-materialize-workspace: obtain sandbox, then Cortex and Memory in parallel (single attempt, 45s fail-fast), then compose model and workflow context.system-nightshift-agent-optimize: obtain sandbox, then Cortex and Memory optimize in settled parallel branches. It replaces the standalone Cortex optimize workflow, declarestool_callsandtool_results(Agent Builder only sendstool_resultsto workflows that declare it), and passes both to Cortex and Memory.system-nightshift-decision-tree-hydrate: restores stored decision trees once per round for the investigation and reinforcement agents.Known limitations
Test plan
cde9c6ad51d(after mergingmainand regenerating the workflow step schemas).cde9c6ad51d. Covers the evidence-only filter (process tools, seeded paths by parameter and bash command, memory reads kept, critique and extraction inputs), extraction skipped without evidence output (empty, reasoning-only, empty-result and seeded-only rounds), the writer budget and compaction, slug idempotency and the truncation boundary, Space isolation, counters, budgets, merge and create-race recovery, OCC retries, source preservation, round-context handoff, telemetry, managed workflow wiring, the Memory transcript renderer (ordering, results, error marking, budgets, fallback, building from hook results), topic entries (new, update, merge, a live memory at the entry's id joining unnamed), the writer writing every entry and skipping empty content, the task as recall key, archived-id reuse, and harmful-memory handling.cde9c6ad51d. This includes real-YAML checks that resume-only steps run at index 0 and not at 1 or 2, that both strict post-execution schemas declare the inputs Agent Builder sends (now includingtool_results), and that both optimize branches receivetool_resultsas an array.5cf304cf8f1: the critique and extraction got every evidence call's result (6 of 19 calls were evidence) and none of the dropped calls; no Cortex, decision-tree or environment-doc content reached any Memory call; extraction and the writer got no final answer. The judgments found the limitations above.3692d5ea962: round 1 26/26; the critique and extraction got only the 7 evidence calls of 20, the writer's prompt carried the budget and it wrote 2,822 characters (no compaction needed).5cf304cf8f1: round 1 26/26; round 2 40/44, where the 4 failures are seed-specific checks after round 1 correctly merged the seed into a new memory. Onf94bcc3629d(after the [Nightshift] Remove model selection from Model Settings #293894 merge): release checks 16/16 (optimize gets the round connector, spend tagged) and HITL 7/7.d548aa2f53c: two rounds, 26/26 each, both optimize branches completed. One round updated two memories in place; the other wrote a new memory through the writer. Every written memory'scontextwas that round's task.a3e6f2fd159: two rounds, 27/27 and 25/25 assertions. The second round updated a memory in place with a recurrence found only in a query result.f74c1e369a6(earlier head): two ordinary turns, 44/44 assertions (memory retrieved and materialized, announced on the first turn only, recalled IDs persisted per round, user message unchanged, hidden context replayed verbatim, no memory IDs in workflow logs, counters persisted); every optimize LLM call carriedfeature_id=significant_events_investigation(9/9); HITL resume 7/7 (materialize and tree hydrate ran only at index 0, optimize and reinforce each ran once with the round's recalled IDs and connector).