Skip to content

feat(kibana): send session id as prompt_cache_key for prompt caching - #122

Merged
vigneshshanmugam merged 1 commit into
elastic:devfrom
flash1293:kibana-prompt-cache-key
Oct 1, 2026
Merged

vigneshshanmugam merged 1 commit into
elastic:devfrom
flash1293:kibana-prompt-cache-key

Conversation

@flash1293

@flash1293 flash1293 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Issue for this PR

Companion to elastic/kibana#294736

Type of change

  • Bug fix
  • New feature
  • Refactor / code improvement
  • Documentation

What does this PR do?

Sends the opencode session id as prompt_cache_key on every request to the Kibana LLM gateway (provider kibana). This is the same thing we already do for OpenRouter.

@ai-sdk/openai-compatible passes unknown providerOptions.kibana.* keys through to the request body, so the field ends up as top-level prompt_cache_key on /internal/elastic_ramen/v1/chat/completions. elastic/kibana#294736 maps it to the inference plugin's sessionId with cacheControl: { type: 'ephemeral', ttl: '5m' }, which EIS uses for prompt caching. All turns of a RAMEN session share one cache scope.

This doesn't re-enable applyCaching for Kibana. Per-message cache_control parts stay off, because the gateway handles caching server-side from the session id.

Backwards compatible: older Kibana versions accept unknown top-level body fields and ignore it.

How did you verify your code works?

  • Unit test in transform.test.ts. There's also an end-to-end llm.test.ts test that checks the HTTP body sent to a mock gateway contains prompt_cache_key === sessionID. It fails without the change.
  • Ran bun dev run locally against a Kibana with the companion PR and EIS. The Phoenix trace for the main-loop chat span showed elastic.cache_control.session_id = ses_…, ttl = 5m.
  • bun typecheck passes; the pre-push hook ran on push.

Checklist

  • I have tested my changes locally
  • I have not included unrelated changes in this PR

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Needs follow-up before merge: the new Kibana prompt_cache_key wiring does not cover the small: true request path.


What is this? | From workflow: PR Review Fork

Give us feedback! React with 🚀 if perfect, 👍 if helpful, 👎 if not.

}

// Kibana gateway forwards prompt_cache_key as the EIS session id to enable prompt caching
if (input.model.providerID === "kibana") {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[MEDIUM] Kibana cache key is only wired on the non-small path

prompt_cache_key is added in ProviderTransform.options, but LLM.stream uses ProviderTransform.smallOptions whenever small: true (for example in title generation). Since smallOptions currently returns {} for Kibana, those requests miss prompt_cache_key, which breaks the PR goal of sending the session cache key on every Kibana request.

Please add the same Kibana cache-key behavior for the small branch as well (or centralize cache-key injection after small/non-small option selection), and add a small: true test to lock this in.

@vigneshshanmugam vigneshshanmugam left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice

@vigneshshanmugam
vigneshshanmugam merged commit 4ff6593 into elastic:dev Oct 1, 2026
13 checks passed
gsoldevila pushed a commit to gsoldevila/kibana that referenced this pull request Oct 2, 2026
…patible endpoint (elastic#294736)

## Summary

The inference plugin already supports `sessionId` + `cacheControl` for
EIS prompt caching (elastic#284774), but the Elastic Console / Elastic Ramen
OpenAI-compatible endpoint
(`/internal/elastic_ramen/v1/chat/completions`) didn't forward them, so
external agents (Elastic Ramen, pi, …) never got prompt caching.

This PR adds that wiring:

- **Session id** is resolved from (first match wins):
1. `prompt_cache_key` body field (standard OpenAI field; used by Elastic
Ramen via `@ai-sdk/openai-compatible`)
2. `x-session-id` header (OpenRouter convention; pi sends this with
`sessionAffinityFormat: "openrouter"`)
  3. `x-session-affinity` header
- When a session id is present, `cacheControl: { type: 'ephemeral', ttl:
'5m' }` is set, matching the Agent Builder main loop.
`prompt_cache_retention: "24h"` maps to the longest EIS TTL (`1h`).
- Usage now includes `prompt_tokens_details.cached_tokens` whenever
inference reports cached tokens, so clients like ai-sdk and pi can show
cache reads.
- Session ids are bounded to 256 chars (validated for the body field;
oversized header values are ignored).
- **Request body limit raised to 20MB** for `/chat/completions` (from
the 1MB `server.maxPayload` default). Long agent conversations with
large tool results or base64 images quickly go over 1MB, and every turn
resends the full history.
- **Message limits raised**: up to 10,000 messages per request (was
1,000) and 2,000 content parts per message (was 100). Each tool call
adds an assistant message and a tool message, so long agent sessions hit
1,000 well before the context window fills. The total payload is still
capped by the 20MB body limit.
- README documents prompt caching and shows a verified pi `models.json`
setup.

Companion Elastic Ramen PR: elastic/elastic-ramen#122

### Verification (local stack with EIS)

Checked through Phoenix traces (`elastic.cache_control.*` span
attributes):
- `prompt_cache_key` in the body → `session_id` + `ttl: 5m` ✅
- `x-session-id` header → `session_id` + `ttl: 5m` ✅
- pi (`openai-completions` + compat config from README) → every turn of
a pi session carries the pi session id ✅. With `PI_CACHE_RETENTION=long`
→ `ttl: 1h` ✅

Body limit: checked live. A ~5MB request returns `200`; a ~21MB request
returns `413 Payload content length greater than maximum allowed:
20971520`.

**Prompt cache hits (Claude Sonnet 5 on EIS):** I sent the same
~422k-token request 3 times with one `prompt_cache_key`:

| Request | Time | `cached_tokens` |
|---|---|---|
| #1 | 8.3s | – (cache written) |
| #2 | 2.1s | 422,347 / 422,349 |
| #3 | 2.0s | 422,347 / 422,349 |

So the session id reaches EIS, EIS caches the prompt, and the new
`prompt_tokens_details.cached_tokens` field reports the hit. Note:
Claude Sonnet 4.5 on EIS doesn't report cached tokens, so no hits show
up for that model.

**Large inputs** (one-word answer hidden in the middle of the input,
sent through the route):

| Model | Prompt tokens | Result |
|---|---|---|
| Claude Sonnet 5 | 422k / 814k | ✅ / ✅ |
| Gemini 2.5 Pro | 388k / 752k | ✅ / ✅ |
| GPT-5.4 | 291k | ✅ |
| Claude Sonnet 4.5 | 124k | ✅ (EIS context window looks like 200k) |

**Long conversations:** a request with 5,001 messages returns `200`.
That was over the old 1,000-message limit.

### Checklist

- [x] [Unit or functional
tests](https://www.elastic.co/guide/en/kibana/master/development-tests.html)
were updated or added to match the most common scenarios
- [x]
[Documentation](https://www.elastic.co/guide/en/kibana/master/development-documentation.html)
was added for features that require explanation or tutorials

## Release Notes

N/A (experimental feature)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants