NVIDIA/ Nemotron
Nemotron 3 Ultra 550B A55B
nvidia/nemotron/3-ultra-550b-a55b
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
Released
2026-06-04
Context window
1M
Max output
128K
Weights
Open weight
Modalities
Input
Text
Output
Text
Identity
- Version
- nemotron-3-ultra-550b-a55b
- Aliases
- Nemotron 3 Ultra 550B A55B
Resources
Providers
| Provider | Provider model ID | Protocol | Context | Max output | Input | Output | Capabilities | Notes | Pricing | Streaming |
|---|---|---|---|---|---|---|---|---|---|---|
| Nvidia | nvidia/nemotron-3-ultra-550b-a55b | custom | 1M | 128K | Text | Text | ReasoningTool callTemperatureOpen weight | - | input $0.5/M · cache read $0.15/M output $2.5/M | Yes |
| Kenari | nemotron-3-ultra-550b-a55b | openai_chat | 1M | 128K | Text | Text | ReasoningTool callTemperatureOpen weight | - | input $0/M output $0/M | Yes |
| UnoRouter | nemotron-3-ultra-550b-a55b:free | openai_chat | 1M | 128K | Text | Text | ReasoningTool callTemperatureOpen weight | - | input $0/M output $0/M | Yes |
| Together AI | nvidia/nemotron-3-ultra-550b-a55b | openai_chat | 1M | 128K | Text | Text | ReasoningTool callTemperatureOpen weight | - | input $0.6/M · cache read $0.2/M output $3.6/M | Yes |
| Vercel AI Gateway | nvidia/nemotron-3-ultra-550b-a55b | openai_chat | 1M | 128K | Text | Text | ReasoningTool callTemperatureOpen weight | - | input $0.6/M · cache read $0.12/M output $2.4/M | Yes |
| OpenRouter | nvidia/nemotron-3-ultra-550b-a55b | openai_chat | 1M | 128K | Text | Text | ReasoningTool callTemperatureOpen weight | Default reasoning: high, Hugging Face: nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | input $0.6/M · cache read $0.2/M output $3.6/M | Yes |
| OpenRouter | nvidia/nemotron-3-ultra-550b-a55b:free | openai_chat | 1M | 128K | Text | Text | ReasoningTool callTemperatureOpen weight | Default reasoning: high, Hugging Face: nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | input $0/M output $0/M | Yes |
| NanoGPT | nvidia/nemotron-3-ultra-550b-a55b | openai_chat | 1M | 128K | Text | Text | ReasoningTool callTemperatureOpen weight | - | input $0.5/M · cache read $0.25/M output $2.5/M | Yes |
| Kilo Gateway | nvidia/nemotron-3-ultra-550b-a55b:free | openai_chat | 1M | 128K | Text | Text | ReasoningTool callTemperatureOpen weight | - | input $0/M output $0/M | Yes |
| Kilo Gateway | nvidia/nemotron-3-ultra-550b-a55b | openai_chat | 1M | 128K | Text | Text | ReasoningTool callTemperatureOpen weight | - | input $0.5/M · cache read $0.1/M output $2.2/M | Yes |
Capabilities
ReasoningTool callTemperatureOpen weight
Official price
input $0.5/M · cache read $0.15/M
output $2.5/M