Back to Registry

Nvidia API Catalog

Comprehensive overview of all nvidia models available through the LLM Kit

Overview

Provider
nvidia
Total Models
96
Last Updated
2026-08-03

Active Speaker Detection

Model ID nvidia/active-speaker-detection
Family
nvidia/active-speaker-detection

Specifications

Context Window: 0 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
video
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

BGE M3

Model ID baai/bge-m3
Family
bge

Specifications

Context Window: 8,192 tokens
Max Output Tokens: 1,024 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

ByteDance-Seed/Seed-OSS-36B-Instruct

Model ID bytedance/seed-oss-36b-instruct
Family
seed

Specifications

Context Window: 262,000 tokens
Max Output Tokens: 262,000 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Cosmos Reason2 8B

Model ID nvidia/cosmos-reason2-8b
Family
nvidia/cosmos-reason2-8b

Specifications

Context Window: 131,072 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text, image, video
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

DeepSeek V4 Flash

Model ID deepseek-ai/deepseek-v4-flash
Family
deepseek-flash

Specifications

Context Window: 1,048,576 tokens
Max Output Tokens: 393,216 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.14
Output
$0.28
Text Tokens - Cached
Input
$0.0
Output
$0.0028

DeepSeek V4 Pro

Model ID deepseek-ai/deepseek-v4-pro
Family
deepseek-thinking

Specifications

Context Window: 1,048,576 tokens
Max Output Tokens: 393,216 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.435
Output
$0.87
Text Tokens - Cached
Input
$0.0
Output
$0.003625

FLUX.1-Kontext-dev

Model ID black-forest-labs/flux_1-kontext-dev
Family
black-forest-labs/flux_1-kontext-dev

Specifications

Context Window: 40,960 tokens
Max Output Tokens: 40,960 tokens

Modalities

Input
text, image
Output
image

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

FLUX.1-dev

Model ID black-forest-labs/flux.1-dev
Family
flux

Specifications

Context Window: 4,096 tokens
Max Output Tokens: 0 tokens

Modalities

Input
text
Output
image

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

FLUX.1-schnell

Model ID black-forest-labs/flux_1-schnell
Family
black-forest-labs/flux_1-schnell

Specifications

Context Window: 77 tokens
Max Output Tokens: 0 tokens

Modalities

Input
text
Output
image

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

FLUX.2 Klein 4B

Model ID black-forest-labs/flux_2-klein-4b
Family
flux

Specifications

Context Window: 40,960 tokens
Max Output Tokens: 40,960 tokens

Modalities

Input
text, image
Output
image

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

GLM-5.2

Model ID z-ai/glm-5.2
Family
glm

Specifications

Context Window: 1,000,000 tokens
Max Output Tokens: 131,072 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

GPT OSS 20B

Model ID openai/gpt-oss-20b
Family
gpt-oss

Specifications

Context Window: 131,072 tokens
Max Output Tokens: 32,768 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

GPT-OSS-120B

Model ID openai/gpt-oss-120b
Family
gpt-oss

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Gemma 2 2b It

Model ID google/gemma-2-2b-it
Family
google/gemma-2-2b-it

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Gemma 3 12B IT

Model ID google/gemma-3-12b-it
Family
gemma

Specifications

Context Window: 131,072 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Gemma 3 4B IT

Model ID google/gemma-3-4b-it
Family
gemma

Specifications

Context Window: 131,072 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Gemma 3n E2b It

Model ID google/gemma-3n-e2b-it
Family
google/gemma-3n-e2b-it

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Gemma 3n E4b It

Model ID google/gemma-3n-e4b-it
Family
google/gemma-3n-e4b-it

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Gemma-4-31B-IT

Model ID google/gemma-4-31b-it
Family
gemma

Specifications

Context Window: 256,000 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text, image, video
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Inkling

Model ID thinkingmachines/inkling
Family
ling

Specifications

Context Window: 1,048,576 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text, audio, image
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Laguna XS 2.1

Model ID poolside/laguna-xs-2.1
Family
laguna

Specifications

Context Window: 262,144 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 3.1 70b Instruct

Model ID meta/llama-3.1-70b-instruct
Family
meta/llama-3.1-70b-instruct

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 3.1 8B Instruct

Model ID meta/llama-3.1-8b-instruct
Family
llama

Specifications

Context Window: 16,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 3.1 Nemotron 70B Instruct

Model ID nvidia/llama-3.1-nemotron-70b-instruct
Family
nemotron

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 3.1 Nemotron Nano 8B v1

Model ID nvidia/llama-3.1-nemotron-nano-8b-v1
Family
nemotron

Specifications

Context Window: 131,072 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 3.1 Nemotron Nano VL 8B v1

Model ID nvidia/llama-3.1-nemotron-nano-vl-8b-v1
Family
nemotron

Specifications

Context Window: 32,768 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 3.1 Nemotron Ultra 253B

Model ID nvidia/llama-3.1-nemotron-ultra-253b-v1
Family
nemotron

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 3.2 11b Vision Instruct

Model ID meta/llama-3.2-11b-vision-instruct
Family
meta/llama-3.2-11b-vision-instruct

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 3.2 1b Instruct

Model ID meta/llama-3.2-1b-instruct
Family
meta/llama-3.2-1b-instruct

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 3.2 3B Instruct

Model ID meta/llama-3.2-3b-instruct
Family
llama

Specifications

Context Window: 32,768 tokens
Max Output Tokens: 32,000 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 3.3 70b Instruct

Model ID meta/llama-3.3-70b-instruct
Family
meta/llama-3.3-70b-instruct

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 3.3 Nemotron Super 49B v1

Model ID nvidia/llama-3.3-nemotron-super-49b-v1
Family
nemotron

Specifications

Context Window: 131,072 tokens
Max Output Tokens: 65,536 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 3.3 Nemotron Super 49B v1.5

Model ID nvidia/llama-3.3-nemotron-super-49b-v1.5
Family
nemotron

Specifications

Context Window: 131,072 tokens
Max Output Tokens: 65,536 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama 4 Maverick 17b 128e Instruct

Model ID meta/llama-4-maverick-17b-128e-instruct
Family
meta/llama-4-maverick-17b-128e-instruct

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama Guard 4 12B

Model ID meta/llama-guard-4-12b
Family
llama

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text, image
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Llama-3.2-90B-Vision-Instruct

Model ID meta/llama-3.2-90b-vision-instruct
Family
llama

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Magistral Small 2506

Model ID mistralai/magistral-small-2506
Family
mistralai/magistral-small-2506

Specifications

Context Window: 32,768 tokens
Max Output Tokens: 32,768 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

MiniMax-M2.7

Model ID minimaxai/minimax-m2.7
Family
minimax

Specifications

Context Window: 204,800 tokens
Max Output Tokens: 131,072 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

MiniMax-M3

Model ID minimaxai/minimax-m3
Family
minimax

Specifications

Context Window: 1,000,000 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text, image, video
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Ministral 3 14B Instruct 2512

Model ID mistralai/ministral-14b-instruct-2512
Family
ministral

Specifications

Context Window: 262,144 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Mistral Large 3 675B Instruct 2512

Model ID mistralai/mistral-large-3-675b-instruct-2512
Family
mistral-large

Specifications

Context Window: 262,144 tokens
Max Output Tokens: 262,144 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Mistral Medium 3

Model ID mistralai/mistral-medium-3-instruct
Family
mistral-medium

Specifications

Context Window: 131,072 tokens
Max Output Tokens: 32,768 tokens

Modalities

Input
text, image
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Mistral Medium 3.5

Model ID mistralai/mistral-medium-3.5-128b
Family
mistral-medium

Specifications

Context Window: 262,144 tokens
Max Output Tokens: 32,768 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Mistral-7B-Instruct-v0.3

Model ID mistralai/mistral-7b-instruct-v0.3
Family
mistralai/mistral-7b-instruct-v0.3

Specifications

Context Window: 65,536 tokens
Max Output Tokens: 65,536 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Mistral: Mixtral 8x22B Instruct

Model ID mistralai/mixtral-8x22b-instruct
Family
mistralai/mixtral-8x22b-instruct

Specifications

Context Window: 65,536 tokens
Max Output Tokens: 13,108 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Mistral: Mixtral 8x7B Instruct

Model ID mistralai/mixtral-8x7b-instruct
Family
mistralai/mixtral-8x7b-instruct

Specifications

Context Window: 32,768 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Nemotron 3 Nano Omni

Model ID nvidia/nemotron-3-nano-omni-30b-a3b-reasoning
Family
nemotron

Specifications

Context Window: 256,000 tokens
Max Output Tokens: 65,536 tokens

Modalities

Input
text, audio, image, video
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Nemotron 3 Super

Model ID nvidia/nemotron-3-super-120b-a12b
Family
nemotron

Specifications

Context Window: 262,144 tokens
Max Output Tokens: 262,144 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.2
Output
$0.8
Text Tokens - Cached
Input
$0.0
Output
$0.0

Nemotron 3 Ultra 550B A55B

Model ID nvidia/nemotron-3-ultra-550b-a55b
Family
nemotron

Specifications

Context Window: 1,000,000 tokens
Max Output Tokens: 65,536 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.5
Output
$2.5
Text Tokens - Cached
Input
$0.0
Output
$0.15

Nemotron Nano 12B v2 VL

Model ID nvidia/nemotron-nano-12b-v2-vl
Family
nemotron

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 128,000 tokens

Modalities

Input
text, image, video
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Phi 4 Multimodal

Model ID microsoft/phi-4-multimodal-instruct
Family
microsoft/phi-4-multimodal-instruct

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Phi-4-Mini

Model ID microsoft/phi-4-mini-instruct
Family
phi

Specifications

Context Window: 131,072 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Qwen Image

Model ID qwen/qwen-image
Family
qwen

Specifications

Context Window: 0 tokens
Max Output Tokens: 0 tokens

Modalities

Input
text, image
Output
image

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Qwen Image Edit

Model ID qwen/qwen-image-edit
Family
qwen

Specifications

Context Window: 0 tokens
Max Output Tokens: 0 tokens

Modalities

Input
text, image
Output
image

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Qwen2.5 Coder 32b Instruct

Model ID qwen/qwen2.5-coder-32b-instruct
Family
qwen/qwen2.5-coder-32b-instruct

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Qwen3 Coder 480B A35B Instruct

Model ID qwen/qwen3-coder-480b-a35b-instruct
Family
qwen

Specifications

Context Window: 262,144 tokens
Max Output Tokens: 66,536 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Qwen3-Next-80B-A3B-Instruct

Model ID qwen/qwen3-next-80b-a3b-instruct
Family
qwen

Specifications

Context Window: 262,144 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Qwen3.5 122B-A10B

Model ID qwen/qwen3.5-122b-a10b
Family
qwen

Specifications

Context Window: 262,144 tokens
Max Output Tokens: 65,536 tokens

Modalities

Input
text, audio, image, video
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Qwen3.5-397B-A17B

Model ID qwen/qwen3.5-397b-a17b
Family
qwen

Specifications

Context Window: 262,144 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Step 3.5 Flash

Model ID stepfun-ai/step-3.5-flash
Family
stepfun-ai/step-3.5-flash

Specifications

Context Window: 256,000 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Step 3.7 Flash

Model ID stepfun-ai/step-3.7-flash
Family
stepfun-ai/step-3.7-flash

Specifications

Context Window: 256,000 tokens
Max Output Tokens: 16,384 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

Whisper Large v3

Model ID openai/whisper-large-v3
Family
whisper

Specifications

Context Window: 0 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
audio
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

bevformer

Model ID nvidia/bevformer
Family
nvidia/bevformer

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
video
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

cosmos-predict1-5b

Model ID nvidia/cosmos-predict1-5b
Family
nvidia/cosmos-predict1-5b

Specifications

Context Window: 0 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text, image, video
Output
video

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

cosmos-transfer1-7b

Model ID nvidia/cosmos-transfer1-7b
Family
nvidia/cosmos-transfer1-7b

Specifications

Context Window: 0 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text, image, video
Output
video

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

cosmos-transfer2.5-2b

Model ID nvidia/cosmos-transfer2_5-2b
Family
nvidia/cosmos-transfer2_5-2b

Specifications

Context Window: 0 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text, image, video
Output
video

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

dracarys-llama-3.1-70b-instruct

Model ID abacusai/dracarys-llama-3.1-70b-instruct
Family
abacusai/dracarys-llama-3.1-70b-instruct

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

esm2-650m

Model ID meta/esm2-650m
Family
meta/esm2-650m

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

esmfold

Model ID meta/esmfold
Family
meta/esmfold

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

gliner-pii

Model ID nvidia/gliner-pii
Family
nvidia/gliner-pii

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

llama-3.1-nemotron-safety-guard-8b-v3

Model ID nvidia/llama-3.1-nemotron-safety-guard-8b-v3
Family
nemotron

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

llama-3_2-nemoretriever-300m-embed-v1

Model ID nvidia/llama-3_2-nemoretriever-300m-embed-v1
Family
nvidia/llama-3_2-nemoretriever-300m-embed-v1

Specifications

Context Window: 32,768 tokens
Max Output Tokens: 2,048 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

llama-nemotron-embed-vl-1b-v2

Model ID nvidia/llama-nemotron-embed-vl-1b-v2
Family
nemotron

Specifications

Context Window: 32,768 tokens
Max Output Tokens: 2,048 tokens

Modalities

Input
text, image
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

llama-nemotron-rerank-vl-1b-v2

Model ID nvidia/llama-nemotron-rerank-vl-1b-v2
Family
nemotron

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text, image
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

magpie-tts-zeroshot

Model ID nvidia/magpie-tts-zeroshot
Family
nvidia/magpie-tts-zeroshot

Specifications

Context Window: 0 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text, audio
Output
audio

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

mistral-nemotron

Model ID mistralai/mistral-nemotron
Family
nemotron

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

mistral-small-4-119b-2603

Model ID mistralai/mistral-small-4-119b-2603
Family
mistralai/mistral-small-4-119b-2603

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text, image
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

nemotron-3-content-safety

Model ID nvidia/nemotron-3-content-safety
Family
nemotron

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

nemotron-3-nano-30b-a3b

Model ID nvidia/nemotron-3-nano-30b-a3b
Family
nemotron

Specifications

Context Window: 131,072 tokens
Max Output Tokens: 131,072 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

nemotron-content-safety-reasoning-4b

Model ID nvidia/nemotron-content-safety-reasoning-4b
Family
nemotron

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Capabilities

Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

nemotron-mini-4b-instruct

Model ID nvidia/nemotron-mini-4b-instruct
Family
nemotron

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

nemotron-voicechat

Model ID nvidia/nemotron-voicechat
Family
nemotron

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text, audio
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

nv-embed-v1

Model ID nvidia/nv-embed-v1
Family
nvidia/nv-embed-v1

Specifications

Context Window: 32,768 tokens
Max Output Tokens: 2,048 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

nv-embedcode-7b-v1

Model ID nvidia/nv-embedcode-7b-v1
Family
nvidia/nv-embedcode-7b-v1

Specifications

Context Window: 32,768 tokens
Max Output Tokens: 2,048 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

nvidia-nemotron-nano-9b-v2

Model ID nvidia/nvidia-nemotron-nano-9b-v2
Family
nemotron

Specifications

Context Window: 131,072 tokens
Max Output Tokens: 131,072 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling Reasoning

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

paligemma

Model ID google/google-paligemma
Family
google/google-paligemma

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text, image
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

rerank-qa-mistral-4b

Model ID nvidia/rerank-qa-mistral-4b
Family
nvidia/rerank-qa-mistral-4b

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

riva-translate-4b-instruct-v1_1

Model ID nvidia/riva-translate-4b-instruct-v1.1
Family
nvidia/riva-translate-4b-instruct-v1.1

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

sarvam-m

Model ID sarvamai/sarvam-m
Family
sarvamai/sarvam-m

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

solar-10.7b-instruct

Model ID upstage/solar-10.7b-instruct
Family
upstage/solar-10.7b-instruct

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text
Output
text

Capabilities

Function calling

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

sparsedrive

Model ID nvidia/sparsedrive
Family
nvidia/sparsedrive

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
video
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

streampetr

Model ID nvidia/streampetr
Family
nvidia/streampetr

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
video
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

studiovoice

Model ID nvidia/studiovoice
Family
nvidia/studiovoice

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 8,192 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

synthetic-video-detector

Model ID nvidia/synthetic-video-detector
Family
nvidia/synthetic-video-detector

Specifications

Context Window: 0 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
video
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

usdcode

Model ID nvidia/usdcode
Family
nvidia/usdcode

Specifications

Context Window: 128,000 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0

usdvalidate

Model ID nvidia/usdvalidate
Family
nvidia/usdvalidate

Specifications

Context Window: 0 tokens
Max Output Tokens: 4,096 tokens

Modalities

Input
text
Output
text

Pricing (per million tokens)

Text Tokens - Standard
Input
$0.0
Output
$0.0
Text Tokens - Cached
Input
$0.0
Output
$0.0