Model information
Released
2026-04-24
Context window
1M
Max output
384K
Model parameters
291B
Weights
Open
Modalities
Input
Text
Output
Text
Identity
- Version
- deepseek-v4-flash
- Knowledge cutoff
- 2025-05
- License
- MIT
- Aliases
- DeepSeek V4 Flash
Resources
Capabilities
Open weightReasoningStructuredTemperatureTool call
Official price
input $0.14/M · output $0.28/M
cache read $0.003/M
Providers
| Provider | Provider model ID | Protocol | Context | Max output | Input | Output | Capabilities | Notes | Pricing | Streaming |
|---|---|---|---|---|---|---|---|---|---|---|
| AnyAPI | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | - | Yes | ||
| CrossModel | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.16/M · output $0.32/M cache read $0.004/M · cache write $0.16/M | Yes | ||
| OpenCode Go | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.14/M · output $0.28/M cache read $0.003/M | Yes | ||
| NanoGPT | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.07/M · output $0.14/M cache read $0.014/M | Yes | ||
| CrofAI | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.12/M · output $0.21/M cache read $0.003/M | Yes | ||
| LLM Gateway | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.076/M · output $0.153/M cache read $0.014/M | Yes | ||
| Kenari | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0/M · output $0/M | Yes | ||
| Kenari | deepseek-v4-flash:free | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0/M · output $0/M | Yes | ||
| OpenCode Zen | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.14/M · output $0.28/M cache read $0.028/M | Yes | ||
| Ambient | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.14/M · output $0.28/M cache read $0.028/M · cache write $0/M | Yes | ||
| UnoRouter | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.063/M · output $0.125/M | Yes | ||
| UnoRouter | deepseek-v4-flash:free | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0/M · output $0/M | Yes | ||
| ZenMux | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.14/M · output $0.28/M cache read $0.003/M | Yes | ||
| EmpirioLabs AI | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.14/M · output $0.28/M cache read $0.14/M | Yes | ||
| Alibaba Token Plan | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0/M · output $0/M cache read $0/M · cache write $0/M | Yes | ||
| Venice AI | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.138/M · output $0.275/M cache read $0.028/M | Yes | ||
| Auriko | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.14/M · output $0.28/M cache read $0.003/M | Yes | ||
| Azure | deepseek-v4-flash | openai_responses | 1M | 384K | ReasoningStructuredTemperature | - | input $0.19/M · output $0.51/M | Yes | ||
| Neuralwatt | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.104/M · output $0.207/M cache read $0.026/M | Yes | ||
| OrcaRouter | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.19/M · output $0.37/M cache read $0.003/M | Yes | ||
| routing.run | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.112/M · output $0.224/M | Yes | ||
| DeepSeek | deepseek-v4-flash | custom | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.14/M · output $0.28/M cache read $0.003/M | Yes | ||
| TensorX | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.15/M · output $0.3/M cache read $0.038/M · cache write $0.188/M | Yes | ||
| Kilo Gateway | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.14/M · output $0.28/M cache read $0.028/M | Yes | ||
| Merge Gateway | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callTemperature | - | input $0.14/M · output $0.28/M cache read $0.003/M | Yes | ||
| Modelis | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.098/M · output $0.197/M | Yes | ||
| Charm Hyper | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.2/M · output $0.4/M cache read $0.04/M | Yes | ||
| Vercel AI Gateway | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.2/M · output $0.4/M cache read $0.04/M | Yes | ||
| Alibaba (China) | deepseek-v4-flash | custom | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.14/M · output $0.28/M cache read $0.003/M | Yes | ||
| NovitaAI | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.14/M · output $0.28/M cache read $0.028/M | Yes | ||
| OpenRouter | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | Reasoning levels: xhigh, high, Default reasoning: high, Hugging Face: deepseek-ai/DeepSeek-V4-Flash | input $0.14/M · output $0.28/M cache read $0.028/M | Yes | ||
| Alibaba Token Plan (China) | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0/M · output $0/M cache read $0/M · cache write $0/M | Yes | ||
| Ollama Cloud | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool call | - | - | Yes | ||
| Cortecs | deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callTemperature | - | input $0.133/M · output $0.266/M cache read $0.003/M | Yes | ||
| HPC-AI | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.14/M · output $0.28/M cache read $0.028/M | Yes | ||
| EBCloud | DeepSeek-V4-Flash | openai_chat | 1M | 384K | ReasoningTool callStructuredTemperature | - | input $0.143/M · output $0.286/M | Yes |