NVIDIA NIM
AI-generatedSummary
Send a conversational chat completion request to a selected NVIDIA NIM AI model using specified conversation messages. Supports optional parameters including max tokens, temperature, top_p, frequency and presence penalties, stop sequences, streaming, and a system prompt prepended to the first user message. The model can be selected from an updated searchable list or by ID with validation format owner/model-name.
Inputs
- Model (required) — The NVIDIA NIM model ID to generate chat responses with, chosen from an expanded searchable list of models including new additions like Llama 3.3, CodeLlama 70B, Mistral variants, Nemotron series, and others, or provided as a string with format owner/model-name.
- Messages (required) — A collection of conversation messages including roles (user or assistant) and contents representing the dialogue history.
- Additional Options — Optional parameters to customize the AI response such as maximum tokens, temperature, top_p, frequency and presence penalties, stop sequences, streaming option, and a system prompt to prepend to the first user message.
Output shape
a list of JSON objects representing the chat completion responses from the selected model, one per input item
Each output JSON includes the full API response for the completion. If an error occurs processing an item and 'continue on fail' is enabled, the output includes an error message for that item.
Examples
Example 1: Generate a chat completion
Select a model (e.g., meta/llama-3.1-8b-instruct), provide a series of user and assistant messages representing conversation history, optionally set parameters such as max tokens and temperature, then run the node to obtain the model's chat response.