Class NvidiaServiceSettings

java.lang.Object
co.elastic.clients.elasticsearch.inference.NvidiaServiceSettings
All Implemented Interfaces:
JsonpSerializable

@JsonpDeserializable public class NvidiaServiceSettings extends Object implements JsonpSerializable
See Also:
  • Field Details

  • Method Details

    • of

    • apiKey

      public final String apiKey()
      Required - A valid API key for your Nvidia endpoint. Can be found in API Keys section of Nvidia account settings.

      API name: api_key

    • url

      @Nullable public final String url()
      The URL of the Nvidia model endpoint. If not provided, the default endpoint URL is used depending on the task type:
      • For text_embedding task - https://integrate.api.nvidia.com/v1/embeddings.
      • For completion and chat_completion tasks - https://integrate.api.nvidia.com/v1/chat/completions.
      • For rerank task - https://ai.api.nvidia.com/v1/retrieval/nvidia/reranking.

      API name: url

    • modelId

      public final String modelId()
      Required - The name of the model to use for the inference task. Refer to the model's documentation for the name if needed. Service has been tested and confirmed to be working with the following models:
      • For text_embedding task - nvidia/llama-3.2-nv-embedqa-1b-v2.
      • For completion and chat_completion tasks - microsoft/phi-3-mini-128k-instruct.
      • For rerank task - nv-rerank-qa-mistral-4b:1. Service doesn't support text_embedding task baai/bge-m3 and nvidia/nvclip models due to them not recognizing the input_type parameter.

      API name: model_id

    • maxInputTokens

      @Nullable public final Integer maxInputTokens()
      For a text_embedding task, the maximum number of tokens per input. Inputs exceeding this value are truncated prior to sending to the Nvidia API.

      API name: max_input_tokens

    • similarity

      @Nullable public final NvidiaSimilarityType similarity()
      For a text_embedding task, the similarity measure. One of cosine, dot_product, l2_norm.

      API name: similarity

    • rateLimit

      @Nullable public final RateLimitSetting rateLimit()
      This setting helps to minimize the number of rate limit errors returned from the Nvidia API. By default, the nvidia service sets the number of requests allowed per minute to 3000.

      API name: rate_limit

    • serialize

      public void serialize(jakarta.json.stream.JsonGenerator generator, JsonpMapper mapper)
      Serialize this object to JSON.
      Specified by:
      serialize in interface JsonpSerializable
    • serializeInternal

      protected void serializeInternal(jakarta.json.stream.JsonGenerator generator, JsonpMapper mapper)
    • toString

      public String toString()
      Overrides:
      toString in class Object
    • rebuild

      Returns:
      New NvidiaServiceSettings.Builder initialized with field values of this instance
    • setupNvidiaServiceSettingsDeserializer

      protected static void setupNvidiaServiceSettingsDeserializer(ObjectDeserializer<NvidiaServiceSettings.Builder> op)