DoublewordDoubleword
Get started

Model Name

deepseek-ai/DeepSeek-V4.1-Flash

DeepSeek V4.1 Flash

  • Type: Generation
  • Capabilities: reasoning
  • Cache read: $0.01 per 1M input tokens (0.02× standard input price). See prompt caching.

Overview

DeepSeek-V4.1-Flash is DeepSeek's September 2026 release: a 552B-parameter MoE with 8B active parameters per input token and 16B per output token, native vision, a 1M-token context window, and a KV cache about a quarter the size of V4-Flash per token. It supports thinking and non-thinking modes with a continuous reasoning-effort setting, DSML tool calling and JSON output.

Reasoning efforts

  • Supported: none, minimal, low, medium, high, xhigh, max

See the reasoning effort guide for request examples.

Pricing

PriorityInput Tokens (per 1M)Cache Read Tokens (per 1M)Output Tokens (per 1M)
Realtime1$0.15$0.01$0.60
Async$0.12$0.01$0.48
Batch (24h)$0.08$0.01$0.30

Playground

Open this model in the Playground.

Footnotes

  1. Realtime availability is limited. Doubleword is primarily a batch API. ↩