Repository navigation
feat(kibana): send session id as prompt_cache_key for prompt caching - #122
Conversation
There was a problem hiding this comment.
Needs follow-up before merge: the new Kibana prompt_cache_key wiring does not cover the small: true request path.
What is this? | From workflow: PR Review Fork
Give us feedback! React with 🚀 if perfect, 👍 if helpful, 👎 if not.
| } | ||
|
|
||
| // Kibana gateway forwards prompt_cache_key as the EIS session id to enable prompt caching | ||
| if (input.model.providerID === "kibana") { |
There was a problem hiding this comment.
[MEDIUM] Kibana cache key is only wired on the non-small path
prompt_cache_key is added in ProviderTransform.options, but LLM.stream uses ProviderTransform.smallOptions whenever small: true (for example in title generation). Since smallOptions currently returns {} for Kibana, those requests miss prompt_cache_key, which breaks the PR goal of sending the session cache key on every Kibana request.
Please add the same Kibana cache-key behavior for the small branch as well (or centralize cache-key injection after small/non-small option selection), and add a small: true test to lock this in.
…patible endpoint (elastic#294736) ## Summary The inference plugin already supports `sessionId` + `cacheControl` for EIS prompt caching (elastic#284774), but the Elastic Console / Elastic Ramen OpenAI-compatible endpoint (`/internal/elastic_ramen/v1/chat/completions`) didn't forward them, so external agents (Elastic Ramen, pi, …) never got prompt caching. This PR adds that wiring: - **Session id** is resolved from (first match wins): 1. `prompt_cache_key` body field (standard OpenAI field; used by Elastic Ramen via `@ai-sdk/openai-compatible`) 2. `x-session-id` header (OpenRouter convention; pi sends this with `sessionAffinityFormat: "openrouter"`) 3. `x-session-affinity` header - When a session id is present, `cacheControl: { type: 'ephemeral', ttl: '5m' }` is set, matching the Agent Builder main loop. `prompt_cache_retention: "24h"` maps to the longest EIS TTL (`1h`). - Usage now includes `prompt_tokens_details.cached_tokens` whenever inference reports cached tokens, so clients like ai-sdk and pi can show cache reads. - Session ids are bounded to 256 chars (validated for the body field; oversized header values are ignored). - **Request body limit raised to 20MB** for `/chat/completions` (from the 1MB `server.maxPayload` default). Long agent conversations with large tool results or base64 images quickly go over 1MB, and every turn resends the full history. - **Message limits raised**: up to 10,000 messages per request (was 1,000) and 2,000 content parts per message (was 100). Each tool call adds an assistant message and a tool message, so long agent sessions hit 1,000 well before the context window fills. The total payload is still capped by the 20MB body limit. - README documents prompt caching and shows a verified pi `models.json` setup. Companion Elastic Ramen PR: elastic/elastic-ramen#122 ### Verification (local stack with EIS) Checked through Phoenix traces (`elastic.cache_control.*` span attributes): - `prompt_cache_key` in the body → `session_id` + `ttl: 5m` ✅ - `x-session-id` header → `session_id` + `ttl: 5m` ✅ - pi (`openai-completions` + compat config from README) → every turn of a pi session carries the pi session id ✅. With `PI_CACHE_RETENTION=long` → `ttl: 1h` ✅ Body limit: checked live. A ~5MB request returns `200`; a ~21MB request returns `413 Payload content length greater than maximum allowed: 20971520`. **Prompt cache hits (Claude Sonnet 5 on EIS):** I sent the same ~422k-token request 3 times with one `prompt_cache_key`: | Request | Time | `cached_tokens` | |---|---|---| | #1 | 8.3s | – (cache written) | | #2 | 2.1s | 422,347 / 422,349 | | #3 | 2.0s | 422,347 / 422,349 | So the session id reaches EIS, EIS caches the prompt, and the new `prompt_tokens_details.cached_tokens` field reports the hit. Note: Claude Sonnet 4.5 on EIS doesn't report cached tokens, so no hits show up for that model. **Large inputs** (one-word answer hidden in the middle of the input, sent through the route): | Model | Prompt tokens | Result | |---|---|---| | Claude Sonnet 5 | 422k / 814k | ✅ / ✅ | | Gemini 2.5 Pro | 388k / 752k | ✅ / ✅ | | GPT-5.4 | 291k | ✅ | | Claude Sonnet 4.5 | 124k | ✅ (EIS context window looks like 200k) | **Long conversations:** a request with 5,001 messages returns `200`. That was over the old 1,000-message limit. ### Checklist - [x] [Unit or functional tests](https://www.elastic.co/guide/en/kibana/master/development-tests.html) were updated or added to match the most common scenarios - [x] [Documentation](https://www.elastic.co/guide/en/kibana/master/development-documentation.html) was added for features that require explanation or tutorials ## Release Notes N/A (experimental feature)
Issue for this PR
Companion to elastic/kibana#294736
Type of change
What does this PR do?
Sends the opencode session id as
prompt_cache_keyon every request to the Kibana LLM gateway (providerkibana). This is the same thing we already do for OpenRouter.@ai-sdk/openai-compatiblepasses unknownproviderOptions.kibana.*keys through to the request body, so the field ends up as top-levelprompt_cache_keyon/internal/elastic_ramen/v1/chat/completions. elastic/kibana#294736 maps it to the inference plugin'ssessionIdwithcacheControl: { type: 'ephemeral', ttl: '5m' }, which EIS uses for prompt caching. All turns of a RAMEN session share one cache scope.This doesn't re-enable
applyCachingfor Kibana. Per-messagecache_controlparts stay off, because the gateway handles caching server-side from the session id.Backwards compatible: older Kibana versions accept unknown top-level body fields and ignore it.
How did you verify your code works?
transform.test.ts. There's also an end-to-endllm.test.tstest that checks the HTTP body sent to a mock gateway containsprompt_cache_key === sessionID. It fails without the change.bun dev runlocally against a Kibana with the companion PR and EIS. The Phoenix trace for the main-loopchatspan showedelastic.cache_control.session_id = ses_…,ttl = 5m.bun typecheckpasses; the pre-push hook ran on push.Checklist