Skip to content

[Inference API] Refactor Contextual AI integration and fix multiple issues - #145700

Merged
Jan-Kazlouski-elastic merged 57 commits into
elastic:mainfrom
Jan-Kazlouski-elastic:feature/contextualai-integration-fixes
Apr 15, 2026
Merged

Jan-Kazlouski-elastic merged 57 commits into
elastic:mainfrom
Jan-Kazlouski-elastic:feature/contextualai-integration-fixes

Conversation

@Jan-Kazlouski-elastic

@Jan-Kazlouski-elastic Jan-Kazlouski-elastic commented Apr 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This pull request refactors the Contextual AI service integration in the inference plugin, fixing critical bugs, removing deprecated settings, and introducing architectural improvements for maintainability and consistency with modern service patterns.

Bugs Fixed

  1. Missing NamedWriteableRegistry registration — ContextualAiRerankServiceSettings and ContextualAiRerankTaskSettings were not registered in InferenceNamedWriteablesProvider, preventing models from being serialized/deserialized over the wire.

  2. HTTP 503 responses not retried — ContextualAiResponseHandler incorrectly treated HTTP 503 as non-retryable. Now correctly retries 503 alongside 500.

  3. Error messages returned HTTP status line instead of body — ContextualAiErrorResponseEntity.fromResponse only returned the status line (e.g., "HTTP/1.1 400 Bad Request"). Replaced with ErrorResponse.fromResponse to read actual response body. ContextualAiErrorResponseEntity deleted entirely.

  4. Duplicate topN fallback logic — ContextualAiRerankRequestEntity contained redundant fallback logic that re-read model task settings. Canonical fallback now lives only in ContextualAiRerankRequest.getTopN().

  5. Overly restrictive type check in doInfer — doInfer rejected models that were ContextualAiModel but not ContextualAiRerankModel. Broadened to accept any ContextualAiModel.

  6. Client-side re-sorting and top-N filtering in response parsing — ContextualAiRerankResponseEntity re-sorted results by relevance score and applied a topN limit client-side. This duplicated server-side behavior. Response parsing now returns results in the order provided by the API without additional filtering.

  7. Input type validation missing — validateInputType was a no-op that accepted any input type. Now properly validates using ServiceUtils.validateInputTypeIsUnspecifiedOrInternal.

Functional Changes

  1. Removed url service setting — ContextualAiRerankServiceSettings no longer accepts/stores the url field. The endpoint URI is hardcoded as DEFAULT_RERANK_URI (https://api.contextual.ai/v1/rerank) in ContextualAiRerankModel. Transport wire backward-compatibility handled via new TransportVersion contextual_ai_url_service_setting_removed with conditional read/write logic: older nodes still send/receive a URL string, which is read and discarded on deserialization, and the default URL is written for older peers on serialization.

  2. Removed hardcoded default instruction — getInstruction() previously always returned "Rerank the given documents based on their relevance to the query." when unset. Now returns null, and the instruction field is omitted from API requests when unset.

  3. Removed document field from response parsing — RankedDocEntry no longer parses the document field from the API response. RankedDoc objects always pass null for the document text, since return_documents was never sent to the ContextualAI API.

  4. Added updateServiceSettings implementation — ContextualAiRerankServiceSettings.updateServiceSettings now supports updating mutable fields (currently only rate_limit). Immutable fields like model_id are preserved from the original settings.

  5. Changed doInfer visibility — doInfer changed from public to protected, consistent with the SenderService base class contract.

Refactoring

  1. Introduced ContextualAiServiceSettings abstract class — Replaces ContextualAiRateLimitServiceSettings interface, encapsulating common model ID and rate-limit settings with shared writeTo/StreamInput, toXContent, fromMap, and updateCommonSettings logic via a CommonSettings record.

  2. Introduced ContextualAiUtils — Centralizes shared transport-version constants (ML_INFERENCE_CONTEXTUAL_AI_ADDED, ML_INFERENCE_CONTEXTUAL_AI_URL_SERVICE_SETTING_REMOVED).

  3. Simplified ContextualAiModel — Removed direct apiKey and rateLimitServiceSettings fields; now delegates to getServiceSettings() and getSecretSettings(). Stores the request URI directly. Added covariant getServiceSettings() / getSecretSettings() overrides returning ContextualAiServiceSettings and DefaultSecretSettings respectively.

  4. Decoupled ContextualAiRerankRequestEntity from ContextualAiRerankModel — Now takes a plain String modelId instead of the model object. Field constants changed from private to public for test assertions.

  5. Simplified ContextualAiRerankRequest — Removed separate instruction constructor parameter; instruction is now retrieved from model.getTaskSettings().getInstruction() at request-creation time. Removed decorateWithAuth helper method and debug logging with try/catch wrapping. Inlined JSON serialization directly.

  6. Simplified ContextualAiRerankResponseEntity — Removed ContextualAiRerankRequest parameter from fromResponse; response parsing no longer depends on the request object.

  7. Inlined createModel helper — Deleted private helper in ContextualAiService.parseRequestConfig, using retrieveModelCreatorFromMapOrThrow(...).createFromMaps(...) directly for consistency with other services.

  8. Added covariant return types — ContextualAiRerankModel.getTaskSettings() returns ContextualAiRerankTaskSettings and getServiceSettings() returns ContextualAiRerankServiceSettings, eliminating scattered casts in callers.

  9. Optimized ContextualAiRerankModel.of() — Returns the same model instance when request task settings are empty or produce no effective change, avoiding unnecessary object creation.

  10. Optimized ContextualAiRerankTaskSettings.of() — Returns the original settings instance when the merged result equals the original, avoiding unnecessary object creation.

Tests

  1. New ContextualAiActionCreatorTests — Integration-style tests using MockWebServer verifying end-to-end request construction (headers, body fields, auth), task settings overrides, request-level topN priority, and error handling for invalid response formats.

  2. New ContextualAiRerankRequestEntityTests — Unit tests verifying XContent serialization with all fields, required-only fields, and null-guard behavior on required parameters.

  3. New ContextualAiRerankRequestTests — Unit tests verifying HTTP request construction (method, URI, headers, auth, body), task settings propagation, and request-level topN override behavior.

  4. New ContextualAiRerankModelTests — Unit tests verifying model construction from maps, ContextualAiRerankModel.of() merge semantics including identity-return optimizations.

  5. New ContextualAiRerankServiceSettingsTests — BWC wire serialization tests, fromMap parsing (required/optional fields, validation errors), updateServiceSettings behavior, and XContent round-tripping.

  6. New ContextualAiRerankTaskSettingsTests — BWC wire serialization tests, fromMap parsing with validation (invalid types, zero/negative topN), of() merge semantics, updatedTaskSettings, isEmpty, and XContent round-tripping.

  7. New ContextualAiRerankResponseEntityTests — Unit tests verifying response parsing for single/multiple items, empty results, floating-point precision, order preservation, unknown fields tolerance, and error cases (missing required fields, malformed JSON).

  8. Updated ContextualAiResponseHandlerTests — Improved tests to verify error body content is included in exception messages, added 503 retryable assertion, extracted shared test fixtures.

  9. New ContextualAiRerankTestFixtures — Shared test constants for inference entity ID, model ID, API key, documents, query, rate limits, and parameterized test values used across all test classes.

Testing

RERANK
ContextualAI API rejects `return_documents`
RQ
{
    "model": "ctxl-rerank-v2-instruct-multilingual-mini",
    "query": "star wars main character",
    "documents": [
        "luke",
        "like",
        "leia",
        "chewy",
        "r2d2",
        "star",
        "wars"
    ],
    "return_documents": true
}
RS
{
    "detail": [
        {
            "msg": "Extra inputs are not permitted",
            "loc": [
                "return_documents"
            ]
        }
    ]
}
Create rerank endpoint (URL present, failure)
RQ
{
    "service": "contextualai",
    "service_settings": {
        "url": "some_url",
        "api_key": "{{contextual-ai-api-key}}",
        "model_id": "ctxl-rerank-v2-instruct-multilingual-mini"
    },
    "task_settings": {
        "instruction": "sort by relevancy",
        "top_n": 2
    }
}
RS
{
    "error": {
        "root_cause": [
            {
                "type": "status_exception",
                "reason": "Configuration contains settings [{url=some_url}] unknown to the [contextualai] service"
            }
        ],
        "type": "status_exception",
        "reason": "Configuration contains settings [{url=some_url}] unknown to the [contextualai] service"
    },
    "status": 400
}
Create rerank endpoint (success)
RQ
{
    "service": "contextualai",
    "service_settings": {
        "api_key": "{{contextual-ai-api-key}}",
        "model_id": "ctxl-rerank-v2-instruct-multilingual-mini"
    },
    "task_settings": {
        "instruction": "sort by relevancy",
        "top_n": 2
    }
}
RS
{
    "inference_id": "contextual-ai-rerank",
    "task_type": "rerank",
    "service": "contextualai",
    "service_settings": {
        "model_id": "ctxl-rerank-v2-instruct-multilingual-mini",
        "rate_limit": {
            "requests_per_minute": 1000
        }
    },
    "task_settings": {
        "top_n": 2,
        "instruction": "sort by relevancy"
    }
}
Perform rerank (all fields)
RQ
{
    "input": [
        "luke",
        "like",
        "leia",
        "chewy",
        "r2d2",
        "star",
        "wars"
    ],
    "query": "star wars main character",
    "top_n": 2,
    "task_settings": {
        "instruction": "classic trilogy only",
        "top_n": 4
    }
}
RS
{
    "rerank": [
        {
            "index": 4,
            "relevance_score": 0.9320585
        },
        {
            "index": 2,
            "relevance_score": 0.9157031
        }
    ]
}
Create rerank endpoint with min fields (success)
RQ
{
    "service": "contextualai",
    "service_settings": {
        "api_key": "{{contextual-ai-api-key}}",
        "model_id": "ctxl-rerank-v2-instruct-multilingual-mini"
    }
}
RS
{
    "inference_id": "contextual-ai-rerank-2",
    "task_type": "rerank",
    "service": "contextualai",
    "service_settings": {
        "model_id": "ctxl-rerank-v2-instruct-multilingual-mini",
        "rate_limit": {
            "requests_per_minute": 1000
        }
    }
}
Perform rerank (minfields)
RQ
{
    "input": [
        "luke",
        "like",
        "leia",
        "chewy",
        "r2d2",
        "star",
        "wars"
    ],
    "query": "star wars main character"
}
RS
{
    "rerank": [
        {
            "index": 4,
            "relevance_score": 0.9514682
        },
        {
            "index": 2,
            "relevance_score": 0.9389061
        },
        {
            "index": 0,
            "relevance_score": 0.9089205
        },
        {
            "index": 5,
            "relevance_score": 0.73254484
        },
        {
            "index": 1,
            "relevance_score": 0.7259128
        },
        {
            "index": 3,
            "relevance_score": 0.6090389
        },
        {
            "index": 6,
            "relevance_score": 0.5869658
        }
    ]
}

References

Register ContextualAiRerankServiceSettings and ContextualAiRerankTaskSettings
in InferenceNamedWriteablesProvider. Require model_id in service settings.
Add ContextualAiUtils for transport version; align writeables with supportsVersion.
Retry HTTP 503; use ErrorResponse::fromResponse for provider error bodies.
Wire return_documents into rerank JSON; simplify top_n serialization.
Enforce non-empty api_key via DefaultSecretSettings path in ContextualAiModel.
Use Elasticsearch logging in ContextualAiRerankRequest; protected doInfer.

Made-with: Cursor
@elasticsearchmachine elasticsearchmachine added external-contributor Pull request authored by a developer outside the Elasticsearch team v9.4.0 labels Apr 3, 2026
.build()
);

configurationMap.put(

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There was a TODO by the engineer that implemented this for giving users ability to specify the URL. Considering there is a such service setting I believe it should be exposed here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It was decided to remove this setting.

config.getTaskType(),
config.getService(),
ConfigurationParseContext.PERSISTENT
ConfigurationParseContext.REQUEST

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Doing this across the services.

}
}

private static ContextualAiModel createModel(

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only used in 1 place. Better to inline.

public static final String NAME = "contextualai";
private static final String SERVICE_NAME = "Contextual AI";

private static final TransportVersion CONTEXTUAL_AI_SERVICE = TransportVersion.fromName("contextual_ai_service");

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Moved version to common place.

int statusCode = result.response().getStatusLine().getStatusCode();
if (statusCode == 500) {
throw new RetryException(true, buildError(SERVER_ERROR, request, result));
} else if (statusCode == 503) {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Retrying both 500s and 503s is standard, added this here as well.

public class ContextualAiResponseHandler extends BaseResponseHandler {

public ContextualAiResponseHandler(String requestType, ResponseParser parseFunction, boolean supportsStreaming) {
super(requestType, parseFunction, ContextualAiErrorResponseEntity::fromResponse, supportsStreaming);

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Original error response didn't have any value


public ContextualAiRerankServiceSettings(URI uri, @Nullable String modelId, @Nullable RateLimitSettings rateLimitSettings) {
this.uri = Objects.requireNonNull(uri);
this.modelId = modelId; // Can be null for REQUEST context

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

model_id is mandatory on ContextualAI side.

@Jan-Kazlouski-elastic

Copy link
Copy Markdown
Contributor Author

Hello @DonalEvans & @jonathan-buttner
Thank you for your comments. They are addressed. PR is ready for another round of review.

@DonalEvans DonalEvans left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, just a couple of small suggestions

* @return the parsed {@link RankedDocsResults}
* @throws IOException if there is an error parsing the response
*/
public static RankedDocsResults fromResponse(HttpResult response) throws IOException {
var parserConfig = XContentParserConfiguration.EMPTY.withDeprecationHandler(LoggingDeprecationHandler.INSTANCE);

try (XContentParser jsonParser = XContentFactory.xContent(XContentType.JSON).createParser(parserConfig, response.body())) {
moveToFirstToken(jsonParser);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm pretty sure this call is no longer needed (if it ever was). If I delete it, all the tests in ContextualAiRerankResponseEntityTests still pass.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Thank you

}
}

private static List<RankedDocsResults.RankedDoc> doParse(XContentParser parser) throws IOException {
private static List<RankedDocsResults.RankedDoc> doParse(XContentParser parser) {
var responseParser = ResponseParser.PARSER;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

responseParser can be inlined, and ResponseParser.PARSER can be replaced with ResponseObject.PARSER and the ResponseParser class removed.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch. Done.

assertThat(
thrownException.getMessage(),
is(Strings.format("unable to parse url [%s]. Reason: Illegal character in path", INVALID_URL))
public void testOf_EmptyMap_RequestOverridesAllValues_AllValuesUpdated() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This test name is a little confusing, since the task settings map is not empty.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Renamed. Now it is clearer. Thanks.

* @return the parsed {@link RankedDocsResults}
* @throws IOException if there is an error parsing the response
*/
public static RankedDocsResults fromResponse(HttpResult response) throws IOException {
var parserConfig = XContentParserConfiguration.EMPTY.withDeprecationHandler(LoggingDeprecationHandler.INSTANCE);

try (XContentParser jsonParser = XContentFactory.xContent(XContentType.JSON).createParser(parserConfig, response.body())) {
moveToFirstToken(jsonParser);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Not your changes but I think we can remove this

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. Thanks.

}
}

private static List<RankedDocsResults.RankedDoc> doParse(XContentParser parser) throws IOException {
private static List<RankedDocsResults.RankedDoc> doParse(XContentParser parser) {
var responseParser = ResponseParser.PARSER;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Also not your changes but I don't think we need the local variable here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

public static boolean supportsContextualAi(TransportVersion version) {
return version.supports(ML_INFERENCE_CONTEXTUAL_AI_ADDED);
}
public static final TransportVersion ML_INFERENCE_CONTEXTUAL_AI_URL_SERVICE_SETTING_REMOVED = TransportVersion.fromName(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How about we drop the ML_ part since we're not in machine learning anymore

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed here and for the initial version as well.

@Jan-Kazlouski-elastic
Jan-Kazlouski-elastic enabled auto-merge (squash) April 14, 2026 10:13
@github-actions

github-actions Bot commented Apr 14, 2026 •

Copy link
Copy Markdown
Contributor

🔍 Preview links for changed docs

⏳ Building and deploying preview... View progress

This comment will be updated with preview links when the build is complete.

@github-actions

Copy link
Copy Markdown
Contributor

ℹ️ Important: Docs version tagging

👋 Thanks for updating the docs! Just a friendly reminder that our docs are now cumulative. This means all 9.x versions are documented on the same page and published off of the main branch, instead of creating separate pages for each minor version.

We use applies_to tags to mark version-specific features and changes.

Expand for a quick overview

When to use applies_to tags:

✅ At the page level to indicate which products/deployments the content applies to (mandatory)
✅ When features change state (e.g. preview, ga) in a specific version
✅ When availability differs across deployments and environments

What NOT to do:

❌ Don't remove or replace information that applies to an older version
❌ Don't add new information that applies to a specific version without an applies_to tag
❌ Don't forget that applies_to tags can be used at the page, section, and inline level

🤔 Need help?

…integration-fixes

# Conflicts:
#	server/src/main/resources/transport/upper_bounds/9.5.csv
…integration-fixes

# Conflicts:
#	server/src/main/resources/transport/upper_bounds/9.5.csv
@Jan-Kazlouski-elastic
Jan-Kazlouski-elastic merged commit fdcf9b2 into elastic:main Apr 15, 2026
40 checks passed
@Jan-Kazlouski-elastic
Jan-Kazlouski-elastic deleted the feature/contextualai-integration-fixes branch April 16, 2026 10:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

>bug external-contributor Pull request authored by a developer outside the Elasticsearch team :SearchOrg/Inference Label for the Search Inference team Team:Search - Inference v9.5.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants