Skip to content

[Elastic Console] Forward session id for prompt caching in OpenAI-compatible endpoint - #294736

Merged
flash1293 merged 3 commits into
elastic:mainfrom
flash1293:elastic-console-session-id
Oct 2, 2026
Merged

flash1293 merged 3 commits into
elastic:mainfrom
flash1293:elastic-console-session-id

Conversation

@flash1293

@flash1293 flash1293 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

The inference plugin already supports sessionId + cacheControl for EIS prompt caching (#284774), but the Elastic Console / Elastic Ramen OpenAI-compatible endpoint (/internal/elastic_ramen/v1/chat/completions) didn't forward them, so external agents (Elastic Ramen, pi, …) never got prompt caching.

This PR adds that wiring:

  • Session id is resolved from (first match wins):
    1. prompt_cache_key body field (standard OpenAI field; used by Elastic Ramen via @ai-sdk/openai-compatible)
    2. x-session-id header (OpenRouter convention; pi sends this with sessionAffinityFormat: "openrouter")
    3. x-session-affinity header
  • When a session id is present, cacheControl: { type: 'ephemeral', ttl: '5m' } is set, matching the Agent Builder main loop. prompt_cache_retention: "24h" maps to the longest EIS TTL (1h).
  • Usage now includes prompt_tokens_details.cached_tokens whenever inference reports cached tokens, so clients like ai-sdk and pi can show cache reads.
  • Session ids are bounded to 256 chars (validated for the body field; oversized header values are ignored).
  • Request body limit raised to 20MB for /chat/completions (from the 1MB server.maxPayload default). Long agent conversations with large tool results or base64 images quickly go over 1MB, and every turn resends the full history.
  • Message limits raised: up to 10,000 messages per request (was 1,000) and 2,000 content parts per message (was 100). Each tool call adds an assistant message and a tool message, so long agent sessions hit 1,000 well before the context window fills. The total payload is still capped by the 20MB body limit.
  • README documents prompt caching and shows a verified pi models.json setup.

Companion Elastic Ramen PR: elastic/elastic-ramen#122

Verification (local stack with EIS)

Checked through Phoenix traces (elastic.cache_control.* span attributes):

  • prompt_cache_key in the body → session_id + ttl: 5m ✅
  • x-session-id header → session_id + ttl: 5m ✅
  • pi (openai-completions + compat config from README) → every turn of a pi session carries the pi session id ✅. With PI_CACHE_RETENTION=long → ttl: 1h ✅

Body limit: checked live. A ~5MB request returns 200; a ~21MB request returns 413 Payload content length greater than maximum allowed: 20971520.

Prompt cache hits (Claude Sonnet 5 on EIS): I sent the same ~422k-token request 3 times with one prompt_cache_key:

Request Time cached_tokens
#1 8.3s – (cache written)
#2 2.1s 422,347 / 422,349
#3 2.0s 422,347 / 422,349

So the session id reaches EIS, EIS caches the prompt, and the new prompt_tokens_details.cached_tokens field reports the hit. Note: Claude Sonnet 4.5 on EIS doesn't report cached tokens, so no hits show up for that model.

Large inputs (one-word answer hidden in the middle of the input, sent through the route):

Model Prompt tokens Result
Claude Sonnet 5 422k / 814k ✅ / ✅
Gemini 2.5 Pro 388k / 752k ✅ / ✅
GPT-5.4 291k ✅
Claude Sonnet 4.5 124k ✅ (EIS context window looks like 200k)

Long conversations: a request with 5,001 messages returns 200. That was over the old 1,000-message limit.

Checklist

Release Notes

N/A (experimental feature)

@flash1293 flash1293 added release_note:skip Skip the PR/issue when compiling release notes backport:skip This PR does not require backporting labels Oct 1, 2026
@flash1293
flash1293 marked this pull request as ready for review October 1, 2026 12:09
@flash1293
flash1293 requested a review from a team as a code owner October 1, 2026 12:09
include_usage: schema.maybe(schema.boolean()),
})
),
// OpenAI prompt caching fields; mapped to the EIS session id / cache control.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity: P3 (Low)

Adding a 256-character schema limit makes chat completions with a longer prompt_cache_key fail request validation, whereas this previously-allowed field was ignored and the new resolver is designed to ignore oversized IDs. Clients that send such a key now lose the entire completion instead of only prompt caching; the README also says these IDs are ignored.

The body object uses { unknowns: 'allow' }, so prompt_cache_key did not previously cause validation failure. The new normalize helper returns undefined for oversized IDs, but the route schema rejects them before it can run.

Generated by Libra. React with 👍 or 👎 to give feedback on this comment.

@kibanamachine

Copy link
Copy Markdown
Contributor

💛 Build succeeded, but was flaky

Failed CI Steps

Metrics [docs]

Unknown metric groups

warm start memory

id before after diff
post forced gc heap baseline - 887739256 +887739256
post forced gc heap delta - -3649314 -3649314
post forced gc heap delta standard deviation - 2328377 +2328377
post forced gc heap target - 884089942 +884089942
tail heap delta - -9064459 -9064459
total +1761443802

Test Failures

  • [job] [logs] FTR Configs #27 / aiops log pattern analysis attaches log pattern analysis table to a dashboard
  • [job] [logs] FTR Configs #4 / discover/ccs_compatible discover search CCS cancel esql mode should show warning and results
  • [job] [logs] FTR Configs #59 / machine learning - stack management jobs import jobs imports jobs

History

@vigneshshanmugam vigneshshanmugam left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

* `x-session-id` matches the OpenRouter convention (used e.g. by the pi agent harness),
* `x-session-affinity` matches the generic OpenAI-compatible affinity header.
*/
const SESSION_ID_HEADERS = ['x-session-id', 'x-session-affinity'] as const;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to pass the sessionAffinityFormat format as well? ?

@flash1293
flash1293 added this pull request to the merge queue Oct 2, 2026
Merged via the queue into elastic:main with commit 038cf22 Oct 2, 2026
42 checks passed
@flash1293
flash1293 deleted the elastic-console-session-id branch October 2, 2026 10:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport:skip This PR does not require backporting release_note:skip Skip the PR/issue when compiling release notes v9.6.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants