Repository navigation
use tabify for regionmap - #14364
Merged
thomasneirynck merged 2 commits intoOct 10, 2017
Merged
use tabify for regionmap#14364
Conversation
Contributor
|
jenkins, test this |
ppisljar
reviewed
Oct 9, 2017
| } else { | ||
| const buckets = response.aggregations[termAggId].buckets; | ||
| const buckets = tableGroup.tables[0].rows; | ||
| results = buckets.map((bucket) => { |
Contributor
There was a problem hiding this comment.
consider ([term, value]) => {
ppisljar
reviewed
Oct 9, 2017
|
|
||
| let results; | ||
| if (!response || !response.aggregations) { | ||
| if (!tableGroup || !tableGroup.tables || !tableGroup.tables.length) { |
Contributor
There was a problem hiding this comment.
add a check that columns have length equal 2 ?
thomasneirynck
added a commit
to thomasneirynck/kibana
that referenced
this pull request
Oct 10, 2017
thomasneirynck
added a commit
to thomasneirynck/kibana
that referenced
this pull request
Oct 10, 2017
Contributor
Author
thomasneirynck
added a commit
that referenced
this pull request
Oct 10, 2017
thomasneirynck
added a commit
that referenced
this pull request
Oct 11, 2017
patrykkopycinski
pushed a commit
to patrykkopycinski/kibana
that referenced
this pull request
May 6, 2026
This was referenced May 8, 2026
Apmats
added a commit
to Apmats/kibana
that referenced
this pull request
May 13, 2026
…lastic#14363) Land the natural-language search API over SML, replacing the transitional bool.should mash with a proper RRF retriever and introducing the trust-boundary split between runtime-imposed scoping and agent-discoverable filters. Decisions per the implementation prompt §6: - §6.1 Retrieval: RRF over BM25 (title^2, description, content via best_fields multi_match) and semantic (unified_semantic). Filters mirrored to each child retriever so RRF can't pull unauthorized docs into the fused top-k. Defaults rank_constant=60, rank_window_size=50. - §6.2 Filter merging: option (2). Service signature splits `scoping?: SmlSearchScoping` (runtime-imposed, unconditional) from `filters?: SmlSearchFilters` (caller refinement). Trust boundary visible at the API layer; merge happens in one place. Sean's per-type id-allowlist shape (elastic#267333) survives unchanged as `SmlSearchScoping`. - §6.3 Agent-supplied dimensions: `types?: string[]`, `tags?: string[]`, each lowering into a `terms` clause. Time-range and free-form payload filters deferred. - §6.4 Compact LLM response: SmlSearchResult drops the `content` blob, full `payload`, dates, and discovery_labels; gains `more_content` (true when the indexed record has non-empty content worth fetching via sml_read). `permissions` retained internally for post-hoc filtering, stripped from the HTTP response. - §6.5 Reference resolution: URI strings only — SML does not dereference. Inline-expansion deferred until eval data justifies it (Reading 1; Sean's "SML shouldn't turn pointers into blobs"). - §6.6 Tool description rewrite with when-to-use guidance and four worked examples (plain query, types+tags filter, connector restriction, wildcard inventory). Runtime scoping is not exposed in the LLM input schema. Fixes the `total` bug: post-permission-filter `results.length` no longer overwrites ES `hits.total.value`. Frontend / HTTP route updates: - POST /sml/_search body accepts `scoping?` and `filters?` (replaces the old `filters` per-type id-allowlist field and `skip_content`). - POST /sml/_autocomplete renames `filters` → `scoping`; no agent-discoverable filters on this route (no LLM in the path). - Frontend SmlService, useSmlSearch, useSmlAutocomplete, usePrefetchSml, queryKeys, and buildSmlScopingFromAgent (formerly buildSmlFiltersFromAgent) all follow the new naming. Tests: - New unit tests for RRF body, filter threading through child retrievers, vocabulary-mismatch over the semantic leg, more_content on/off, compact result mapping, agent-filter builder, and tool wrapper scoping/filters mapping. - FTR + Scout integration tests updated to drop content/skip_content assertions and lock in the no-content invariant. Out of scope (deferred per epic): reranker, kneedle, per-type retriever profiles, HyDE/query expansion, reference inline-expansion, telemetry instrumentation (elastic#14366), recrawl/versioning (elastic#14367). Open coordination points (not blocking this PR): - Inference model on unified_semantic still relies on defaults (Kathleen Mar 19: pin Jina V5). Belongs to A. - Kathleen review on A's PR proposes splitting unified_semantic into per-field semantic_text fields. If A reverts unified, the semantic leg in this PR fans out from 1 to 2-3 child retrievers — small retriever-only change. - Pierre's configuration_override-aware tool handler context (elastic#267333 review) — wrapper currently reads agentConfiguration from the tool handler context, which `runner.ts` already resolves with overrides applied via `resolveConfiguration`. Note in code. Stacked on top of apmats/sml-autocomplete-followup (elastic#14364). Once autocomplete merges, this PR rebases to main. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Apmats
added a commit
that referenced
this pull request
Jun 5, 2026
…earch (#269277) # [Agent Builder] [SML] Search API: hybrid retrieval, autocomplete, schema Closes search-team#14362, search-team#14363, search-team#14364 --- ## Summary Implements the SML search and autocomplete API for the `agent_context_layer` plugin. Three issues combined into this PR: schema + mappings (#14362), hybrid search API (#14363), autocomplete decoupling (#14364). --- ## Schema The index (`.chat-sml-data`) uses `dynamic: strict` mappings. Full mapping: ```json { "dynamic": "strict", "properties": { "id": { "type": "keyword" }, "type": { "type": "keyword" }, "title": { "type": "text", "fields": { "semantic": { "type": "semantic_text" } } }, "origin_id": { "type": "keyword" }, "origin": { "type": "object", "properties": { "uri": { "type": "keyword" } } }, "content": { "type": "text", "fields": { "semantic": { "type": "semantic_text" } } }, "description": { "type": "text", "fields": { "semantic": { "type": "semantic_text" } } }, "tags": { "type": "keyword", "normalizer": "lowercase" }, "discovery_labels": { "type": "nested", "properties": { "value": { "type": "search_as_you_type" }, "kind": { "type": "keyword" } } }, "references": { "type": "object", "properties": { "uri": { "type": "keyword" } } }, "extended_attrs": { "type": "flattened" }, "user_id": { "type": "keyword" }, "created_at": { "type": "date" }, "updated_at": { "type": "date" }, "spaces": { "type": "keyword" }, "permissions": { "type": "keyword" } } } ``` Key schema decisions (see RFC for full reasoning): - **`title`, `description`, `content`**: `text` + `.semantic` multi-field (`semantic_text`). Separate vectors per field for independent recall in RRF. - **`extended_attrs`**: `flattened` - every leaf is keyword-searchable without per-type schema changes. SML treats it opaquely; type producers own its shape. Sub-field filtering not exposed in v1. - **`discovery_labels`**: `nested { value: search_as_you_type, kind: keyword }`. The indexer auto-prepends `{value: title, kind: 'title'}` and `{value: type, kind: 'type'}` to every record. Producers add arbitrary extra entries (taglines, nicknames, categories). The `kind` field lets the UI decide how to render a matched label. - **`references`**: `object { uri: keyword }` rather than `keyword[]`. Functionally identical today; the object shape leaves room for a relationship `kind` sub-field without a future migration. - **`origin`**: Self-describing URI object (`origin.uri = "${type}://${originId}"`) constructed at index time. `origin_id` is retained internally for `deleteByQuery` and constraints filters. - **`tags`**: `keyword` with `lowercase` normalizer for case-insensitive matching. Agent-discoverable filter facet via `filters.tags`. --- ## Routes ### `POST /internal/agent_context_layer/sml/_search` Hybrid retrieval via ES|QL `FORK + FUSE`. Two branches - BM25 (`MATCH` over `title`, `description`, `content`) and semantic (`MATCH` over `.semantic` multi-fields) - fused with RRF. Space filter applied as `WHERE MV_CONTAINS(spaces, ?)` before the fork. Empty/wildcard queries fall back to a plain sorted scan. **Request:** ```json { "query": "cpu usage over time", "size": 10, "constraints": { "connector": { "ids": ["abc123"] } }, "filters": { "types": ["dashboard"], "tags": ["production"] }, "fields": ["content", "tags", "references"] } ``` **Response:** ```json { "results": [ { "id": "dashboard:xyz789:uuid", "type": "dashboard", "origin": { "uri": "dashboard://xyz789" }, "title": "Infrastructure CPU Overview", "description": "Shows CPU usage trends across production hosts", "content": "...", "tags": ["production", "infrastructure"], "references": [{ "uri": "index://metrics-*" }] } ] } ``` **`fields[]` opt-in.** If omitted, returns `id`, `type`, `title`, `origin`, `description`. If provided, `id`, `type`, `title`, `origin` are always included; `description` and any other listed fields (`content`, `tags`, `references`, `spaces`, `permissions`) are included only if requested. `spaces` and `permissions` are always fetched server-side (space scoping and RBAC) but stripped from the response unless listed. Unknown values are rejected at the route schema level (`schema.oneOf` literals). **No score.** RRF scores are rank-fusion heuristics with no calibrated semantic meaning. Exposing them risks the LLM drawing false confidence. Can be added as an opt-in field later if needed for diagnostics. ### `POST /internal/agent_context_layer/sml/_autocomplete` Nested `multi_match bool_prefix` with `operator: and` over `discovery_labels.value` and its search-as-you-type subfields. `inner_hits` surfaces which specific label matched and its `kind`. `search_as_you_type bool_prefix` chosen over `edge_ngram` for tighter matching semantics (all-but-last tokens are exact; only the trailing partial is a prefix). Server-side highlights are broken due to ES bug [#53744](elastic/elasticsearch#53744) (open since 2020); provisions are in place for when it is fixed. **Request:** ```json { "query": "git", "size": 5, "constraints": { "connector": { "ids": ["abc123"] } }, "filters": { "types": ["connector"] } } ``` **Response:** ```json { "results": [ { "id": "connector:abc123:uuid", "type": "connector", "origin": { "uri": "connector://abc123" }, "title": "GitHub Connector", "matched_discovery_labels": [ { "value": "github", "kind": "tagline" }, { "value": "GitHub Connector", "kind": "title" } ] } ] } ``` --- ## Key Design Decisions ### `constraints` vs `filters` as separate parameters Two separate parameters on both routes and the service interface - `constraints: SmlSearchConstraints` and `filters: SmlSearchFilters`: - `constraints` - runtime-imposed per-type id-allowlists. Applied by the call wrapper from the caller's context (agent SO `connector_ids`, future allowed-indices, RBAC). Not in the LLM tool input schema. The agent can't see or modify this. - `filters` - agent-discoverable `{ types?, tags? }` refinements. In the tool input schema. Agent can set these, but they can only narrow `constraints`, never widen them. This makes the trust boundary visible at the API layer. `AgentContextLayerPluginStart.search()` owns the merge - call wrappers pass two clearly separated inputs. Adding a new runtime constraint doesn't require touching every wrapper. The shape deliberately supports only id-allowlists. Cross-type constraints compose with OR. More complex constraints must be pre-computed into a flat id list or expressed as a separate named parameter. ### Separate routes for autocomplete and full retrieval Before these PRs, autocomplete and search shared a single route. The @ menu and the LLM `sml_search` tool have different latency budgets, different result shapes, and different query mechanisms. Two dedicated routes: - `_autocomplete` - no content, no scores, no description. Tuned for @ menu rendering. - `_search` - full result, `fields[]` opt-in. Tuned for LLM consumption. ### ES|QL FORK + FUSE instead of DSL RRF retriever The original POC used `retriever.rrf` DSL. Replaced with ES|QL `FORK + FUSE` per team direction. Two branches: BM25 (`MATCH` OR across text fields) + semantic (`MATCH` OR across `.semantic` multi-fields), fused via `FUSE`. ### No inline resolution of `references` `references` are returned as URI strings. SML never dereferences them. Callers that need to follow a reference make an explicit lookup themselves. Auto-resolving would work against the goal of low token utilization - pulling full content blobs speculatively rather than on demand. > **Future:** A dedicated `sml_read` lookup API (tracked separately) is the intended home for reference resolution. Once that exists, multi-hop traversal (depth > 1) becomes possible caller-side: fetch a result, follow a `references` URI via `sml_read`, repeat. How many hops, at what content budget, and whether to expand eagerly or lazily are left as caller-side decisions - SML does not need to encode a traversal policy. ### No query language, no expressive text search operators `query` is a single opaque string injected verbatim into the FORK template. Nothing is parsed from it. The only structured search surface is `filters.types` and `filters.tags`. A larger API surface means larger tool descriptions and more exploratory LLM tool calls; additions need to justify themselves against that cost. --- ## Post-RFC Changes (from team review) The RFC was approved and then the following changes were made based on feedback before this PR: | Change | Before | After | |---|---|---| | Score | Returned in results | Dropped - RRF score has no calibrated meaning | | Content opt-in | `skip_content: boolean` | `fields: string[]` - explicit field selection, forward-compatible | | Origin identifier | `origin_id: string` | `origin: { uri: string }` - self-describing, no need to reconstruct from `type` + raw id | | Type-specific blob | `payload` | `extended_attrs` - clearer that it's supplementary producer-owned data, not the primary carrier | | Runtime allowlist param | `scoping` / `SmlSearchScoping` | `constraints` / `SmlSearchConstraints` | | Query engine | DSL `retriever.rrf` | ES|QL `FORK + FUSE` | --- ## What's deliberately not in v1 - `extended_attrs` sub-field filtering not exposed in `SmlSearchFilters` (ES `flattened` supports it; no agreed filter shape yet) - No `/sml/_facets` endpoint (no consistent tag vocabulary across producers yet) - No query language operators (`type:dashboard`, `+must -exclude`) - no evidence they'd help over baseline - No inline reference resolution - progressive disclosure story not settled - No reranking or per-field boost tuning - intentionally naive equal-weight baseline until eval data exists --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: kibanamachine <42973632+kibanamachine@users.noreply.github.com>
patrykkopycinski
pushed a commit
to patrykkopycinski/kibana
that referenced
this pull request
Aug 5, 2026
…earch (elastic#269277) # [Agent Builder] [SML] Search API: hybrid retrieval, autocomplete, schema Closes search-team#14362, search-team#14363, search-team#14364 --- ## Summary Implements the SML search and autocomplete API for the `agent_context_layer` plugin. Three issues combined into this PR: schema + mappings (elastic#14362), hybrid search API (elastic#14363), autocomplete decoupling (elastic#14364). --- ## Schema The index (`.chat-sml-data`) uses `dynamic: strict` mappings. Full mapping: ```json { "dynamic": "strict", "properties": { "id": { "type": "keyword" }, "type": { "type": "keyword" }, "title": { "type": "text", "fields": { "semantic": { "type": "semantic_text" } } }, "origin_id": { "type": "keyword" }, "origin": { "type": "object", "properties": { "uri": { "type": "keyword" } } }, "content": { "type": "text", "fields": { "semantic": { "type": "semantic_text" } } }, "description": { "type": "text", "fields": { "semantic": { "type": "semantic_text" } } }, "tags": { "type": "keyword", "normalizer": "lowercase" }, "discovery_labels": { "type": "nested", "properties": { "value": { "type": "search_as_you_type" }, "kind": { "type": "keyword" } } }, "references": { "type": "object", "properties": { "uri": { "type": "keyword" } } }, "extended_attrs": { "type": "flattened" }, "user_id": { "type": "keyword" }, "created_at": { "type": "date" }, "updated_at": { "type": "date" }, "spaces": { "type": "keyword" }, "permissions": { "type": "keyword" } } } ``` Key schema decisions (see RFC for full reasoning): - **`title`, `description`, `content`**: `text` + `.semantic` multi-field (`semantic_text`). Separate vectors per field for independent recall in RRF. - **`extended_attrs`**: `flattened` - every leaf is keyword-searchable without per-type schema changes. SML treats it opaquely; type producers own its shape. Sub-field filtering not exposed in v1. - **`discovery_labels`**: `nested { value: search_as_you_type, kind: keyword }`. The indexer auto-prepends `{value: title, kind: 'title'}` and `{value: type, kind: 'type'}` to every record. Producers add arbitrary extra entries (taglines, nicknames, categories). The `kind` field lets the UI decide how to render a matched label. - **`references`**: `object { uri: keyword }` rather than `keyword[]`. Functionally identical today; the object shape leaves room for a relationship `kind` sub-field without a future migration. - **`origin`**: Self-describing URI object (`origin.uri = "${type}://${originId}"`) constructed at index time. `origin_id` is retained internally for `deleteByQuery` and constraints filters. - **`tags`**: `keyword` with `lowercase` normalizer for case-insensitive matching. Agent-discoverable filter facet via `filters.tags`. --- ## Routes ### `POST /internal/agent_context_layer/sml/_search` Hybrid retrieval via ES|QL `FORK + FUSE`. Two branches - BM25 (`MATCH` over `title`, `description`, `content`) and semantic (`MATCH` over `.semantic` multi-fields) - fused with RRF. Space filter applied as `WHERE MV_CONTAINS(spaces, ?)` before the fork. Empty/wildcard queries fall back to a plain sorted scan. **Request:** ```json { "query": "cpu usage over time", "size": 10, "constraints": { "connector": { "ids": ["abc123"] } }, "filters": { "types": ["dashboard"], "tags": ["production"] }, "fields": ["content", "tags", "references"] } ``` **Response:** ```json { "results": [ { "id": "dashboard:xyz789:uuid", "type": "dashboard", "origin": { "uri": "dashboard://xyz789" }, "title": "Infrastructure CPU Overview", "description": "Shows CPU usage trends across production hosts", "content": "...", "tags": ["production", "infrastructure"], "references": [{ "uri": "index://metrics-*" }] } ] } ``` **`fields[]` opt-in.** If omitted, returns `id`, `type`, `title`, `origin`, `description`. If provided, `id`, `type`, `title`, `origin` are always included; `description` and any other listed fields (`content`, `tags`, `references`, `spaces`, `permissions`) are included only if requested. `spaces` and `permissions` are always fetched server-side (space scoping and RBAC) but stripped from the response unless listed. Unknown values are rejected at the route schema level (`schema.oneOf` literals). **No score.** RRF scores are rank-fusion heuristics with no calibrated semantic meaning. Exposing them risks the LLM drawing false confidence. Can be added as an opt-in field later if needed for diagnostics. ### `POST /internal/agent_context_layer/sml/_autocomplete` Nested `multi_match bool_prefix` with `operator: and` over `discovery_labels.value` and its search-as-you-type subfields. `inner_hits` surfaces which specific label matched and its `kind`. `search_as_you_type bool_prefix` chosen over `edge_ngram` for tighter matching semantics (all-but-last tokens are exact; only the trailing partial is a prefix). Server-side highlights are broken due to ES bug [elastic#53744](elastic/elasticsearch#53744) (open since 2020); provisions are in place for when it is fixed. **Request:** ```json { "query": "git", "size": 5, "constraints": { "connector": { "ids": ["abc123"] } }, "filters": { "types": ["connector"] } } ``` **Response:** ```json { "results": [ { "id": "connector:abc123:uuid", "type": "connector", "origin": { "uri": "connector://abc123" }, "title": "GitHub Connector", "matched_discovery_labels": [ { "value": "github", "kind": "tagline" }, { "value": "GitHub Connector", "kind": "title" } ] } ] } ``` --- ## Key Design Decisions ### `constraints` vs `filters` as separate parameters Two separate parameters on both routes and the service interface - `constraints: SmlSearchConstraints` and `filters: SmlSearchFilters`: - `constraints` - runtime-imposed per-type id-allowlists. Applied by the call wrapper from the caller's context (agent SO `connector_ids`, future allowed-indices, RBAC). Not in the LLM tool input schema. The agent can't see or modify this. - `filters` - agent-discoverable `{ types?, tags? }` refinements. In the tool input schema. Agent can set these, but they can only narrow `constraints`, never widen them. This makes the trust boundary visible at the API layer. `AgentContextLayerPluginStart.search()` owns the merge - call wrappers pass two clearly separated inputs. Adding a new runtime constraint doesn't require touching every wrapper. The shape deliberately supports only id-allowlists. Cross-type constraints compose with OR. More complex constraints must be pre-computed into a flat id list or expressed as a separate named parameter. ### Separate routes for autocomplete and full retrieval Before these PRs, autocomplete and search shared a single route. The @ menu and the LLM `sml_search` tool have different latency budgets, different result shapes, and different query mechanisms. Two dedicated routes: - `_autocomplete` - no content, no scores, no description. Tuned for @ menu rendering. - `_search` - full result, `fields[]` opt-in. Tuned for LLM consumption. ### ES|QL FORK + FUSE instead of DSL RRF retriever The original POC used `retriever.rrf` DSL. Replaced with ES|QL `FORK + FUSE` per team direction. Two branches: BM25 (`MATCH` OR across text fields) + semantic (`MATCH` OR across `.semantic` multi-fields), fused via `FUSE`. ### No inline resolution of `references` `references` are returned as URI strings. SML never dereferences them. Callers that need to follow a reference make an explicit lookup themselves. Auto-resolving would work against the goal of low token utilization - pulling full content blobs speculatively rather than on demand. > **Future:** A dedicated `sml_read` lookup API (tracked separately) is the intended home for reference resolution. Once that exists, multi-hop traversal (depth > 1) becomes possible caller-side: fetch a result, follow a `references` URI via `sml_read`, repeat. How many hops, at what content budget, and whether to expand eagerly or lazily are left as caller-side decisions - SML does not need to encode a traversal policy. ### No query language, no expressive text search operators `query` is a single opaque string injected verbatim into the FORK template. Nothing is parsed from it. The only structured search surface is `filters.types` and `filters.tags`. A larger API surface means larger tool descriptions and more exploratory LLM tool calls; additions need to justify themselves against that cost. --- ## Post-RFC Changes (from team review) The RFC was approved and then the following changes were made based on feedback before this PR: | Change | Before | After | |---|---|---| | Score | Returned in results | Dropped - RRF score has no calibrated meaning | | Content opt-in | `skip_content: boolean` | `fields: string[]` - explicit field selection, forward-compatible | | Origin identifier | `origin_id: string` | `origin: { uri: string }` - self-describing, no need to reconstruct from `type` + raw id | | Type-specific blob | `payload` | `extended_attrs` - clearer that it's supplementary producer-owned data, not the primary carrier | | Runtime allowlist param | `scoping` / `SmlSearchScoping` | `constraints` / `SmlSearchConstraints` | | Query engine | DSL `retriever.rrf` | ES|QL `FORK + FUSE` | --- ## What's deliberately not in v1 - `extended_attrs` sub-field filtering not exposed in `SmlSearchFilters` (ES `flattened` supports it; no agreed filter shape yet) - No `/sml/_facets` endpoint (no consistent tag vocabulary across producers yet) - No query language operators (`type:dashboard`, `+must -exclude`) - no evidence they'd help over baseline - No inline reference resolution - progressive disclosure story not settled - No reranking or per-field boost tuning - intentionally naive equal-weight baseline until eval data exists --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: kibanamachine <42973632+kibanamachine@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Use the tabify responseHandler in regionmaps
Similar to #14266.