Repository navigation
[Elastic Console] Forward session id for prompt caching in OpenAI-compatible endpoint - #294736
Conversation
| include_usage: schema.maybe(schema.boolean()), | ||
| }) | ||
| ), | ||
| // OpenAI prompt caching fields; mapped to the EIS session id / cache control. |
There was a problem hiding this comment.
Severity: P3 (Low)
Adding a 256-character schema limit makes chat completions with a longer prompt_cache_key fail request validation, whereas this previously-allowed field was ignored and the new resolver is designed to ignore oversized IDs. Clients that send such a key now lose the entire completion instead of only prompt caching; the README also says these IDs are ignored.
The body object uses { unknowns: 'allow' }, so prompt_cache_key did not previously cause validation failure. The new normalize helper returns undefined for oversized IDs, but the route schema rejects them before it can run.
Generated by Libra. React with 👍 or 👎 to give feedback on this comment.
💛 Build succeeded, but was flaky
Failed CI StepsMetrics [docs]Unknown metric groupswarm start memory
Test Failures
History
|
| * `x-session-id` matches the OpenRouter convention (used e.g. by the pi agent harness), | ||
| * `x-session-affinity` matches the generic OpenAI-compatible affinity header. | ||
| */ | ||
| const SESSION_ID_HEADERS = ['x-session-id', 'x-session-affinity'] as const; |
There was a problem hiding this comment.
Do we want to pass the sessionAffinityFormat format as well? ?
Summary
The inference plugin already supports
sessionId+cacheControlfor EIS prompt caching (#284774), but the Elastic Console / Elastic Ramen OpenAI-compatible endpoint (/internal/elastic_ramen/v1/chat/completions) didn't forward them, so external agents (Elastic Ramen, pi, …) never got prompt caching.This PR adds that wiring:
prompt_cache_keybody field (standard OpenAI field; used by Elastic Ramen via@ai-sdk/openai-compatible)x-session-idheader (OpenRouter convention; pi sends this withsessionAffinityFormat: "openrouter")x-session-affinityheadercacheControl: { type: 'ephemeral', ttl: '5m' }is set, matching the Agent Builder main loop.prompt_cache_retention: "24h"maps to the longest EIS TTL (1h).prompt_tokens_details.cached_tokenswhenever inference reports cached tokens, so clients like ai-sdk and pi can show cache reads./chat/completions(from the 1MBserver.maxPayloaddefault). Long agent conversations with large tool results or base64 images quickly go over 1MB, and every turn resends the full history.models.jsonsetup.Companion Elastic Ramen PR: elastic/elastic-ramen#122
Verification (local stack with EIS)
Checked through Phoenix traces (
elastic.cache_control.*span attributes):prompt_cache_keyin the body →session_id+ttl: 5m✅x-session-idheader →session_id+ttl: 5m✅openai-completions+ compat config from README) → every turn of a pi session carries the pi session id ✅. WithPI_CACHE_RETENTION=long→ttl: 1h✅Body limit: checked live. A ~5MB request returns
200; a ~21MB request returns413 Payload content length greater than maximum allowed: 20971520.Prompt cache hits (Claude Sonnet 5 on EIS): I sent the same ~422k-token request 3 times with one
prompt_cache_key:cached_tokensSo the session id reaches EIS, EIS caches the prompt, and the new
prompt_tokens_details.cached_tokensfield reports the hit. Note: Claude Sonnet 4.5 on EIS doesn't report cached tokens, so no hits show up for that model.Large inputs (one-word answer hidden in the middle of the input, sent through the route):
Long conversations: a request with 5,001 messages returns
200. That was over the old 1,000-message limit.Checklist
Release Notes
N/A (experimental feature)