Skip to content

[ESQL] Add coordinator-only caching for external source metadata - #145300

Merged
costin merged 3 commits into
elastic:mainfrom
costin:esql/schema-inference-caching
Apr 1, 2026
Merged

costin merged 3 commits into
elastic:mainfrom
costin:esql/schema-inference-caching

Conversation

@costin

@costin costin commented Mar 31, 2026 •

Copy link
Copy Markdown
Member

Summary

Introduce in-memory caching for schema inference and file listing
results on the coordinator node, reducing repeated S3/storage API
calls for external data sources.

Two-cache design

Schema and listing use separate caches with different TTLs,
budgets, key strategies, and sharing models because they have
fundamentally different invalidation characteristics:

Schema cache Listing cache
What changes Only when the file is rewritten (rare) Whenever files are added/removed (frequent)
Budget 20% of total 80% of total
TTL 5 min (schema is stable; re-reading a Parquet footer is expensive) 30 sec (listing must stay fresh to pick up new files)
Key includes credentials? No — schema is determined by file content, not who reads it Yes — different credentials may see different S3 buckets/paths
Invalidation signal mtime of the anchor file embedded in the key Hard TTL only (no mtime signal for a directory listing)

Why 80/20 instead of 50/50: The split reflects the current
first_file_wins-only schema caching, where a single schema entry
(~2.5 KB for 30 columns) is cached per glob pattern while a
compressed listing entry stores per-file data for potentially
thousands of files (~32–52 KB for 1,000 files after compaction).
Even compressed, a listing entry is 13–20× larger than a single
schema entry.

This ratio only holds because strict and union_by_name
reconciliation bypasses the schema cache entirely today — it
reads all files' schemas on every query. If per-file schema caching
were added for reconciliation mode (1,000 files × 2.5 KB ≈ 2.5 MB
vs 52 KB listing), the schema side would dominate and the budget
split would need to be revisited or made adaptive.

Keeping the caches separate means a stale listing (30s) doesn't
force a schema re-read (which would waste another Parquet footer
fetch), and a long-lived schema entry doesn't pin a stale listing
in memory.

Caching details

  • ExternalSourceCacheService manages both caches on the
    coordinator; total budget defaults to 0.4% of heap
    (esql.source.cache.size)
  • SchemaCacheKey = anchor URI + mtime + endpoint + region +
    format-affecting params (delimiter, encoding, etc.) — credentials
    are excluded so schema is shared across users accessing the same
    file
  • ListingCacheKey = scheme + bucket + glob + endpoint + region +
    MurmurHash3-128 of credential values — credentials are hashed
    (never stored) to isolate listings per user
  • NameId-safe SchemaCacheEntry reconstructs fresh
    ReferenceAttribute instances per query so cached attributes
    never share NameIds across queries
  • Only the first_file_wins schema resolution mode uses caching;
    strict and union_by_name reconciliation bypasses the schema
    cache since it needs all files' schemas

File list refactoring

  • FileList SPI interface with UNRESOLVED/EMPTY sentinels and
    optional fileSchemaInfo() for schema reconciliation
  • datasources.glob package with GlobExpander as sole public facade
  • GenericFileList, DictionaryFileList (~52 B/file), HiveFileList
    (~32 B/file) are package-private — consumers use only FileList
  • FileListCompactor orchestrates compaction without coupling between
    the data classes (no HiveFileList ↔ DictionaryFileList dependency)
  • Renamed FileSet → GenericFileList, fileSet() → fileList()

Developed with AI-assisted tooling

@costin
costin requested a review from bpintea March 31, 2026 09:16
@costin
costin enabled auto-merge (squash) March 31, 2026 09:17
@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Hi @costin, I've created a changelog YAML for you.

@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Pinging @elastic/es-analytical-engine (Team:Analytics)

@elasticsearchmachine elasticsearchmachine added the Team:Analytics Meta label for analytical engine team (ESQL/Aggs/Geo) label Mar 31, 2026
@github-actions

github-actions Bot commented Mar 31, 2026 •

Copy link
Copy Markdown
Contributor

🔍 Preview links for changed docs

⏳ Building and deploying preview... View progress

This comment will be updated with preview links when the build is complete.

@github-actions

Copy link
Copy Markdown
Contributor

ℹ️ Important: Docs version tagging

👋 Thanks for updating the docs! Just a friendly reminder that our docs are now cumulative. This means all 9.x versions are documented on the same page and published off of the main branch, instead of creating separate pages for each minor version.

We use applies_to tags to mark version-specific features and changes.

Expand for a quick overview

When to use applies_to tags:

✅ At the page level to indicate which products/deployments the content applies to (mandatory)
✅ When features change state (e.g. preview, ga) in a specific version
✅ When availability differs across deployments and environments

What NOT to do:

❌ Don't remove or replace information that applies to an older version
❌ Don't add new information that applies to a specific version without an applies_to tag
❌ Don't forget that applies_to tags can be used at the page, section, and inline level

🤔 Need help?

@costin
costin force-pushed the esql/schema-inference-caching branch 3 times, most recently from 1f92ddd to 81dcedf Compare March 31, 2026 12:04

@bpintea bpintea left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖-assisted review.

costin added 3 commits April 1, 2026 11:56
Introduce in-memory caching for schema inference and file listing
results on the coordinator node, reducing repeated S3/storage API
calls for external data sources.

Key components:
- ExternalSourceCacheService with independent schema (20% budget,
  5m TTL) and listing (80% budget, 30s TTL) caches
- NameId-safe SchemaCacheEntry that reconstructs fresh Attributes
  per query to prevent cross-query interference
- Credential-isolated ListingCacheKey using MurmurHash3-128
- DictionaryFileList (~52B/file) and HiveFileList (~32B/file)
  compact representations, always used for multi-file sources
- FileList SPI interface replacing concrete FileSet dependency
- Renamed FileSet -> GenericFileList, fileSet() -> fileList()

Developed with AI-assisted tooling
Consolidate GlobExpander, GenericFileList, DictionaryFileList, and
HiveFileList into the datasources.glob package. All concrete FileList
implementations are now package-private; consumers depend only on
the FileList interface. GlobExpander is the sole public entry point
for creating and compacting file lists.

Extract compaction logic from HiveFileList.from/DictionaryFileList.from
into a package-private FileListCompactor, eliminating the dependency
between the two data classes. Both are now pure data holders.

UNRESOLVED and EMPTY sentinels are inline anonymous implementations
on the FileList interface.
@costin
costin force-pushed the esql/schema-inference-caching branch from 3c0d68d to 6371cd6 Compare April 1, 2026 10:50
@costin
costin disabled auto-merge April 1, 2026 16:18
@costin
costin merged commit 1ae9e2c into elastic:main Apr 1, 2026
32 of 35 checks passed
@costin
costin deleted the esql/schema-inference-caching branch April 1, 2026 17:20
mromaios pushed a commit to mromaios/elasticsearch that referenced this pull request Apr 9, 2026
…stic#145300)

Introduce in-memory caching for schema inference and file listing
results on the coordinator node, reducing repeated S3/storage API
calls for external data sources.

Schema and listing use separate caches with different TTLs,
budgets, key strategies, and sharing models because they have
fundamentally different invalidation characteristics. The current
split is 20/80 and reflects the current first_file_wins-only schema 
caching, where a single schema entry
(~2.5 KB for 30 columns) is cached per glob pattern while a
compressed listing entry stores per-file data for potentially
thousands of files (~32–52 KB for 1,000 files after compaction).
Even compressed, a listing entry is 13–20× larger than a single
schema entry.

## Caching details
ExternalSourceCacheService manages both caches on the
coordinator; total budget defaults to 0.4% of heap
(esql.source.cache.size)
SchemaCacheKey = anchor URI + mtime + endpoint + region +
format-affecting params (delimiter, encoding, etc.) — credentials
are excluded so schema is shared across users accessing the same
file
ListingCacheKey = scheme + bucket + glob + endpoint + region +
MurmurHash3-128 of credential values — credentials are hashed
(never stored) to isolate listings per user
NameId-safe SchemaCacheEntry reconstructs fresh
ReferenceAttribute instances per query so cached attributes
never share NameIds across queries
Only the first_file_wins schema resolution mode uses caching;
strict and union_by_name reconciliation bypasses the schema
cache since it needs all files' schemas

Keeping the caches separate means a stale listing (30s) doesn't
force a schema re-read (which would waste another Parquet footer
fetch), and a long-lived schema entry doesn't pin a stale listing
in memory.

Developed with AI-assisted tooling
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

:Analytics/ES|QL AKA ESQL >enhancement ES|QL|DS ES|QL datasources Team:Analytics Meta label for analytical engine team (ESQL/Aggs/Geo) v9.4.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants