Model Name
deepseek-ai/DeepSeek-V4.1-FlashDeepSeek V4.1 Flash
- Type: Generation
- Capabilities:
reasoning - Cache read: $0.01 per 1M input tokens (0.02× standard input price). See prompt caching.
Overview
DeepSeek-V4.1-Flash is DeepSeek's September 2026 release: a 552B-parameter MoE with 8B active parameters per input token and 16B per output token, native vision, a 1M-token context window, and a KV cache about a quarter the size of V4-Flash per token. It supports thinking and non-thinking modes with a continuous reasoning-effort setting, DSML tool calling and JSON output.
Reasoning efforts
- Supported:
none,minimal,low,medium,high,xhigh,max
See the reasoning effort guide for request examples.
Pricing
| Priority | Input Tokens (per 1M) | Cache Read Tokens (per 1M) | Output Tokens (per 1M) |
|---|---|---|---|
| Realtime1 | $0.15 | $0.01 | $0.60 |
| Async | $0.12 | $0.01 | $0.48 |
| Batch (24h) | $0.08 | $0.01 | $0.30 |
Playground
Open this model in the Playground.
Footnotes
-
Realtime availability is limited. Doubleword is primarily a batch API. ↩