DeepSeek/ DeepSeek
DeepSeek V4 Flash
deepseek/deepseek/v4-flash
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
Released
2026-04-24
Context window
1M
Max output
384K
Weights
Open weight
Modalities
Input
Text
Output
Text
Identity
- Version
- deepseek-v4-flash
- Knowledge cutoff
- 2025-05
- Aliases
- DeepSeek V4 Flash
Resources
Providers
| Provider | Provider model ID | Protocol | Context | Max output | Input | Output | Capabilities | Notes | Pricing | Streaming |
|---|---|---|---|---|---|---|---|---|---|---|
| AnyAPI | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | - | Yes |
| CrossModel | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0.16/M · cache read $0.004/M output $0.32/M · cache write $0.16/M | Yes |
| OpenCode Go | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.14/M · cache read $0.0028/M output $0.28/M | Yes |
| NanoGPT | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0.07/M · cache read $0.014/M output $0.14/M | Yes |
| CrofAI | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.12/M · cache read $0.003/M output $0.21/M | Yes |
| LLM Gateway | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.076/M · cache read $0.014/M output $0.153/M | Yes |
| Kenari | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0/M output $0/M | Yes |
| Kenari | deepseek-v4-flash:free | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0/M output $0/M | Yes |
| OpenCode Zen | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.14/M · cache read $0.028/M output $0.28/M | Yes |
| Ambient | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0.14/M · cache read $0.028/M output $0.28/M · cache write $0/M | Yes |
| UnoRouter | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0.0625/M output $0.125/M | Yes |
| UnoRouter | deepseek-v4-flash:free | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0/M output $0/M | Yes |
| ZenMux | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.14/M · cache read $0.0028/M output $0.28/M | Yes |
| EmpirioLabs AI | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0.14/M · cache read $0.14/M output $0.28/M | Yes |
| Alibaba Token Plan | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0/M · cache read $0/M output $0/M · cache write $0/M | Yes |
| Venice AI | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.138/M · cache read $0.028/M output $0.275/M | Yes |
| Auriko | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.14/M · cache read $0.0028/M output $0.28/M | Yes |
| Azure | deepseek-v4-flash | openai_responses | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0.19/M output $0.51/M | Yes |
| Neuralwatt | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0.104/M · cache read $0.026/M output $0.207/M | Yes |
| OrcaRouter | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.19/M · cache read $0.0028/M output $0.37/M | Yes |
| routing.run | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.112/M output $0.224/M | Yes |
| DeepSeek | deepseek-v4-flash | custom | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.14/M · cache read $0.0028/M output $0.28/M | Yes |
| TensorX | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0.15/M · cache read $0.0375/M output $0.3/M · cache write $0.1875/M | Yes |
| Kilo Gateway | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0.14/M · cache read $0.028/M output $0.28/M | Yes |
| Merge Gateway | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.14/M · cache read $0.0028/M output $0.28/M | Yes |
| Modelis | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0.0983/M output $0.1966/M | Yes |
| Charm Hyper | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0.2/M · cache read $0.04/M output $0.4/M | Yes |
| Vercel AI Gateway | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | input $0.2/M · cache read $0.04/M output $0.4/M | Yes |
| Alibaba (China) | deepseek-v4-flash | custom | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.14/M · cache read $0.0028/M output $0.28/M | Yes |
| NovitaAI | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.14/M · cache read $0.028/M output $0.28/M | Yes |
| OpenRouter | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output, Default reasoning: high, Hugging Face: deepseek-ai/DeepSeek-V4-Flash | input $0.14/M · cache read $0.028/M output $0.28/M | Yes |
| Alibaba Token Plan (China) | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0/M · cache read $0/M output $0/M · cache write $0/M | Yes |
| Ollama Cloud | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | - | - | Yes |
| Cortecs | deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.133/M · cache read $0.0028/M output $0.266/M | Yes |
| HPC-AI | deepseek/deepseek-v4-flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.14/M · cache read $0.028/M output $0.28/M | Yes |
| EBCloud | DeepSeek-V4-Flash | openai_chat | 1M | 384K | Text | Text | ReasoningTool callStructuredTemperatureOpen weight | Interleaved output | input $0.143/M output $0.2857/M | Yes |
Capabilities
ReasoningTool callStructuredTemperatureOpen weight
Official price
input $0.14/M · cache read $0.0028/M
output $0.28/M