Skip to content

[Inference API] Task settings are ignored for all Amazon Bedrock and certain Google Vertex chat completions #148792

Description

@DonalEvans

When creating a chat_completion endpoint using the amazonbedrock or googlevertexai services, it is possible to specify a task_settings field containing some settings that are applicable to chat_completion. For Amazon Bedrock, the relevant settings are max_new_tokens, temperature, top_p and top_k. For Google Vertex, the relevant setting is max_tokens.

When task settings are specified at endpoint creation time, they should be used for all requests using that endpoint unless explicitly overridden in the request. However, when an AmazonBedrockChatCompletionRequestEntity is created in AmazonBedrockChatCompletionEntityFactory, the values used for max_new_tokens, temperature and top_p are taken only from the request, with the task settings associated with the model being ignored, and the top_k setting being omitted entirely, since it can't be specified in a request (see the code here).

Similarly, for a GoogleVertexAiUnifiedChatCompletionRequestEntity (used when the provider service settings is omitted or set to google), the task settings are ignored, meaning that the max_tokens task setting is not used (see the code here).

To fix this bug, the values configured in the task settings should be used if the corresponding values in the request are null.

For Amazon Bedrock chat completions, this bug is present in 9.4 and onward (introduced in #139411, which added chat_completion for Amazon Bedrock).
For Google Vertex, this bug is present in 9.2 and onward (introduced in #134080, which added the max_tokens task setting).

Activity

  1. elasticsearchmachine commented on May 11, 2026

    @elasticsearchmachine
    Collaborator

    Pinging @elastic/search-inference-team (Team:Search - Inference)

  2. lost-particles commented on May 18, 2026

    @lost-particles
    Contributor

    Opened #149268 with a fix. For the Bedrock side I added a topK field on the unified entity and routed it through the same additional-fields channel the non-unified path uses, so the wiring is consistent. For Vertex it's just passing the configured max_tokens through and falling back to it when the per-request value is null. Maintainers, let me know if anything needs adjusting.

  3. DonalEvans commented on May 19, 2026

    @DonalEvans
    ContributorAuthor

    The PR to fix this issue on main should be able to be backported to 9.4 cleanly, but for 9.3, only the Google Vertex changes can be backported, since Amazon Bedrock didn't support chat completion in 9.3. I'll take care of any backport merge conflicts once the main PR is merged.

  4. added a commit that references this issue on May 19, 2026
    b10718b
  5. added a commit that references this issue on May 20, 2026
    7d225d4
  6. added a commit that references this issue on May 20, 2026
    c869663
  7. added 2 commits that reference this issue on May 22, 2026
    5f3ace5
    21ce42e
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions