- Providers
- Nvidia
First party
Nvidia
Nvidia is a First party provider with 31 models across 31 offerings.
Access
- API base URL
- https://integrate.api.nvidia.com/v1
- SDK
- @ai-sdk/openai-compatible
- Environment variables
- NVIDIA_API_KEY
Protocols
Openai chatCustom
Organizations
| Organization | Families | Models |
|---|---|---|
| Alibaba | 1 | 3 |
| 1 | 4 | |
| Meta | 1 | 1 |
| Mistral AI | 1 | 1 |
| Moonshot AI | 1 | 1 |
| NVIDIA | 1 | 15 |
| OpenAI | 2 | 3 |
| POPoolside | 1 | 1 |
| TMThinking Machines | 1 | 1 |
| Zhipu AI | 1 | 1 |
Models
31 of 31 model versions
| Model | Organization | Family | Provider model ID | Context | Max output | Input | Output | Reasoning | Tool call | Structured | Temperature | Weights | Pricing | Released | Updated | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Inkling thinkingmachines/ling/inkling | TMThinking Machines | Ling | thinkingmachines/inkling Custom · Streaming | 1M | 1M | TextImageAudio | Text | Yes | Yes | -No | Yes | Open weight | input $0/M output $0/M | 2026-07-15 | 2026-07-17 | |
| Laguna XS 2.1 poolside/laguna/xs-2-1 | POPoolside | Laguna | poolside/laguna-xs-2.1 Custom · Streaming | 262K | 33K | Text | Text | Yes | Yes | -No | Yes | Open weight | input $0/M output $0/M | 2026-07-02 | 2026-07-02 | |
| GLM-5.2 zhipuai/glm/5-2 | Zhipu AI | GLM | z-ai/glm-5.2 Openai chat · Streaming | 1M | 131K | Text | Text | Yes | Yes | Yes | Yes | Open weight | input $0/M output $0/M | 2026-06-13 | 2026-07-22 | |
| Nemotron 3 Ultra 550B A55B nvidia/nemotron/3-ultra-550b-a55b | NVIDIA | Nemotron | nvidia/nemotron-3-ultra-550b-a55b Custom · Streaming | 1M | 128K | Text | Text | Yes | Yes | -No | Yes | Open weight | input $0.5/M · cache read $0.15/M output $2.5/M | 2026-06-04 | 2026-06-04 | |
| Nemotron 3 Nano Omni 30B A3B Reasoning nvidia/nemotron/3-nano-omni-30b-a3b-reasoning | NVIDIA | Nemotron | nvidia/nemotron-3-nano-omni-30b-a3b-reasoning Custom · Streaming | 256K | 66K | TextImageVideoAudio | Text | Yes | Yes | -No | Yes | Open weight | input $0/M output $0/M | 2026-04-28 | 2026-04-28 | |
| Kimi K2.6 moonshotai/kimi/k2-6 | Moonshot AI | Kimi | moonshotai/kimi-k2.6 Custom · Streaming | 262K | 262K | TextImageVideo | Text | Yes | Yes | Yes | Yes | Open weight | input $0/M output $0/M | 2026-04-21 | 2026-07-22 | |
| Nemotron 3 Content Safety nvidia/nemotron/3-content-safety | NVIDIA | Nemotron | nvidia/nemotron-3-content-safety Custom · Streaming | 128K | 4.1K | Text | Text | -No | -No | -No | -No | Open weight | input $0/M output $0/M | 2026-04-16 | 2026-04-16 | |
| Gemma 4 31B IT google/gemma/4-31b-it | Gemma | google/gemma-4-31b-it Custom · Streaming | 262K | 33K | TextImage | Text | Yes | Yes | Yes | Yes | Open weight | input $0/M output $0/M | 2026-04-02 | 2026-04-30 | ||
| Llama Nemotron Rerank VL 1B v2 nvidia/nemotron/llama-nemotron-rerank-vl-1b-v2 | NVIDIA | Nemotron | nvidia/llama-nemotron-rerank-vl-1b-v2 Custom · Streaming | 128K | 4.1K | TextImage | Text | -No | -No | -No | -No | Open weight | input $0/M output $0/M | 2026-03-31 | 2026-03-31 | |
| Nemotron VoiceChat nvidia/nemotron/voicechat | NVIDIA | Nemotron | nvidia/nemotron-voicechat Custom · Streaming | 128K | 8.2K | TextAudio | Text | -No | Yes | -No | Yes | Open weight | input $0/M output $0/M | 2026-03-16 | 2026-03-16 | |
| Nemotron 3 Super 120B A12B nvidia/nemotron/3-super-120b-a12b | NVIDIA | Nemotron | nvidia/nemotron-3-super-120b-a12b Custom · Streaming | 262K | 262K | Text | Text | Yes | Yes | -No | Yes | Open weight | input $0.2/M output $0.8/M | 2026-03-11 | 2026-03-12 | |
| Qwen3.5 122B-A10B alibaba/qwen/qwen3-5-122b-a10b | Alibaba | Qwen | qwen/qwen3.5-122b-a10b Openai chat · Streaming | 262K | 66K | TextImageVideoAudio | Text | Yes | Yes | Yes | Yes | Open weight | input $0/M output $0/M | 2026-02-23 | 2026-03-18 | |
| Qwen3.5 397B-A17B alibaba/qwen/qwen3-5-397b-a17b | Alibaba | Qwen | qwen/qwen3.5-397b-a17b Openai chat · Streaming | 262K | 66K | TextImageVideoAudio | Text | Yes | Yes | Yes | Yes | Open weight | input $0/M output $0/M | 2026-02-15 | 2026-06-11 | |
| Llama Nemotron Embed VL 1B v2 nvidia/nemotron/llama-nemotron-embed-vl-1b-v2 | NVIDIA | Nemotron | nvidia/llama-nemotron-embed-vl-1b-v2 Custom · Streaming | 33K | 2K | TextImage | Text | -No | -No | -No | -No | Open weight | input $0/M output $0/M | 2026-02-10 | 2026-02-10 | |
| Nemotron Content Safety Reasoning 4B nvidia/nemotron/content-safety-reasoning-4b | NVIDIA | Nemotron | nvidia/nemotron-content-safety-reasoning-4b Custom · Streaming | 128K | 4.1K | Text | Text | Yes | -No | -No | -No | Open weight | input $0/M output $0/M | 2026-01-22 | 2026-01-22 | |
| Nemotron 3 Nano 30B A3B nvidia/nemotron/3-nano-30b-a3b | NVIDIA | Nemotron | nvidia/nemotron-3-nano-30b-a3b Custom · Streaming | 262K | 262K | Text | Text | Yes | Yes | -No | Yes | Open weight | input $0/M output $0/M | 2025-12-15 | 2025-12-15 | |
| Llama 3.1 Nemotron Safety Guard 8B v3 nvidia/nemotron/llama-3-1-nemotron-safety-guard-8b-v3 | NVIDIA | Nemotron | nvidia/llama-3.1-nemotron-safety-guard-8b-v3 Custom · Streaming | 128K | 4.1K | Text | Text | -No | -No | -No | -No | Open weight | input $0/M output $0/M | 2025-10-28 | 2025-10-28 | |
| Nemotron Nano 12B v2 VL nvidia/nemotron/nano-12b-v2-vl | NVIDIA | Nemotron | nvidia/nemotron-nano-12b-v2-vl Custom · Streaming | 128K | 128K | TextImageVideo | Text | Yes | Yes | -No | Yes | Open weight | input $0/M output $0/M | 2025-10-28 | 2025-10-28 | |
| Qwen3-Next 80B-A3B Instruct alibaba/qwen/qwen3-next-80b-a3b-instruct | Alibaba | Qwen | qwen/qwen3-next-80b-a3b-instruct Openai chat · Streaming | 131K | 33K | Text | Text | -No | Yes | Yes | Yes | Open weight | input $0/M output $0/M | 2025-09 | 2026-07-22 | |
| GPT OSS 120B openai/gpt/oss-120b | OpenAI | GPT | openai/gpt-oss-120b Custom · Streaming | 131K | 33K | Text | Text | Yes | Yes | Yes | Yes | Open weight | input $0/M output $0/M | 2025-08-05 | 2026-07-22 | |
| GPT OSS 20B openai/gpt/oss-20b | OpenAI | GPT | openai/gpt-oss-20b Custom · Streaming | 131K | 33K | Text | Text | Yes | Yes | Yes | Yes | Open weight | input $0/M output $0/M | 2025-08-05 | 2026-03-01 | |
| Llama 3.3 Nemotron Super 49B v1.5 nvidia/nemotron/llama-3-3-nemotron-super-49b-v1-5 | NVIDIA | Nemotron | nvidia/llama-3.3-nemotron-super-49b-v1.5 Custom · Streaming | 131K | 131K | Text | Text | Yes | Yes | -No | Yes | Open weight | input $0/M output $0/M | 2025-07-25 | 2025-07-25 | |
| Gemma 3n 4B google/gemma/3n-e4b-it | Gemma | google/gemma-3n-e4b-it Openai chat · Streaming | 33K | - | Text | Text | -No | -No | Yes | Yes | Open weight | input $0/M output $0/M | 2025-05-20 | 2025-06-03 | ||
| Llama 3.1 Nemotron 70B Instruct nvidia/nemotron/llama-3-1-nemotron-70b-instruct | NVIDIA | Nemotron | nvidia/llama-3.1-nemotron-70b-instruct Custom · Streaming | 128K | 8.2K | Text | Text | -No | Yes | -No | Yes | Open weight | input $0/M output $0/M | 2025-04-15 | 2025-04-15 | |
| Llama 3.3 Nemotron Super 49B v1 nvidia/nemotron/llama-3-3-nemotron-super-49b-v1 | NVIDIA | Nemotron | nvidia/llama-3.3-nemotron-super-49b-v1 Custom · Streaming | 131K | 131K | Text | Text | Yes | Yes | -No | Yes | Open weight | input $0/M output $0/M | 2025-04-07 | 2025-08-08 | |
| Gemma 3 12B google/gemma/3-12b-it | Gemma | google/gemma-3-12b-it Openai chat · Streaming | 131K | 16K | TextImage | Text | -No | Yes | Yes | Yes | Open weight | input $0/M output $0/M | 2025-03-13 | 2025-03-13 | ||
| Gemma 3 4B google/gemma/3-4b-it | Gemma | google/gemma-3-4b-it Openai chat · Streaming | 131K | 16K | TextImage | Text | -No | -No | Yes | Yes | Open weight | input $0/M output $0/M | 2025-03-13 | 2025-03-13 | ||
| Llama-3.3-70B-Instruct meta/llama/3-3-70b-instruct | Meta | Llama | meta/llama-3.3-70b-instruct Custom · Streaming | 128K | 4.1K | Text | Text | -No | Yes | Yes | Yes | Open weight | input $0/M output $0/M | 2024-12-06 | 2026-07-22 | |
| Whisper 3 Large openai/whisper/large-v3 | OpenAI | Whisper | openai/whisper-large-v3 Custom · Streaming | 448 | 4.1K | Audio | Text | -No | -No | -No | -No | Open weight | input $0/M output $0/M | 2024-10-01 | 2026-03-17 | |
| Nemotron Mini 4B Instruct nvidia/nemotron/mini-4b-instruct | NVIDIA | Nemotron | nvidia/nemotron-mini-4b-instruct Custom · Streaming | 128K | 8.2K | Text | Text | -No | Yes | -No | Yes | Open weight | input $0/M output $0/M | 2024-08-21 | 2024-08-26 | |
| Mixtral 8x22B Instruct mistral/mixtral/8x22b-instruct | Mistral AI | Mixtral | mistralai/mixtral-8x22b-instruct Openai chat · Streaming | 66K | - | TextDocument | Text | -No | Yes | Yes | Yes | Open weight | input $0/M output $0/M | 2024-04-17 | 2024-04-17 |