Class NvidiaServiceSettings
java.lang.Object
co.elastic.clients.elasticsearch.inference.NvidiaServiceSettings
- All Implemented Interfaces:
JsonpSerializable
- See Also:
-
Nested Class Summary
Nested Classes -
Field Summary
FieldsModifier and TypeFieldDescriptionstatic final JsonpDeserializer<NvidiaServiceSettings>Json deserializer forNvidiaServiceSettings -
Method Summary
Modifier and TypeMethodDescriptionfinal StringapiKey()Required - A valid API key for your Nvidia endpoint.final IntegerFor atext_embeddingtask, the maximum number of tokens per input.final StringmodelId()Required - The name of the model to use for the inference task.static NvidiaServiceSettingsfinal RateLimitSettingThis setting helps to minimize the number of rate limit errors returned from the Nvidia API.rebuild()voidserialize(jakarta.json.stream.JsonGenerator generator, JsonpMapper mapper) Serialize this object to JSON.protected voidserializeInternal(jakarta.json.stream.JsonGenerator generator, JsonpMapper mapper) protected static voidfinal NvidiaSimilarityTypeFor atext_embeddingtask, the similarity measure.toString()final Stringurl()The URL of the Nvidia model endpoint.
-
Field Details
-
_DESERIALIZER
Json deserializer forNvidiaServiceSettings
-
-
Method Details
-
of
public static NvidiaServiceSettings of(Function<NvidiaServiceSettings.Builder, ObjectBuilder<NvidiaServiceSettings>> fn) -
apiKey
Required - A valid API key for your Nvidia endpoint. Can be found inAPI Keyssection of Nvidia account settings.API name:
api_key -
url
The URL of the Nvidia model endpoint. If not provided, the default endpoint URL is used depending on the task type:- For
text_embeddingtask -https://integrate.api.nvidia.com/v1/embeddings. - For
completionandchat_completiontasks -https://integrate.api.nvidia.com/v1/chat/completions. - For
reranktask -https://ai.api.nvidia.com/v1/retrieval/nvidia/reranking.
API name:
url - For
-
modelId
Required - The name of the model to use for the inference task. Refer to the model's documentation for the name if needed. Service has been tested and confirmed to be working with the following models:- For
text_embeddingtask -nvidia/llama-3.2-nv-embedqa-1b-v2. - For
completionandchat_completiontasks -microsoft/phi-3-mini-128k-instruct. - For
reranktask -nv-rerank-qa-mistral-4b:1. Service doesn't supporttext_embeddingtaskbaai/bge-m3andnvidia/nvclipmodels due to them not recognizing theinput_typeparameter.
API name:
model_id - For
-
maxInputTokens
For atext_embeddingtask, the maximum number of tokens per input. Inputs exceeding this value are truncated prior to sending to the Nvidia API.API name:
max_input_tokens -
similarity
For atext_embeddingtask, the similarity measure. One of cosine, dot_product, l2_norm.API name:
similarity -
rateLimit
This setting helps to minimize the number of rate limit errors returned from the Nvidia API. By default, thenvidiaservice sets the number of requests allowed per minute to 3000.API name:
rate_limit -
serialize
Serialize this object to JSON.- Specified by:
serializein interfaceJsonpSerializable
-
serializeInternal
-
toString
-
rebuild
- Returns:
- New
NvidiaServiceSettings.Builderinitialized with field values of this instance
-
setupNvidiaServiceSettingsDeserializer
protected static void setupNvidiaServiceSettingsDeserializer(ObjectDeserializer<NvidiaServiceSettings.Builder> op)
-