DoublewordDoubleword

Model Name

Qwen/Qwen3-VL-30B-A3B-Instruct-FP8

Qwen3 VL 30B A3B Instruct

  • Type: Generation
  • Capabilities: vision
  • Cache read: $0.04 per 1M input tokens (0.25× standard input price). See prompt caching.

Overview

Meet Qwen3-VL-30B, the smaller model of the Qwen3-VL family, delivering performance similar to GPT-4.1-mini and Claude Sonnet 4. This highly capable mid-size model is suited for tasks that are constrained or require high token volumes. Excels at reasoning, coding, and structured output generation.

Best for:

  • Production workloads requiring strong performance without frontier model costs
  • Complex reasoning tasks
  • Code generation

Pricing

PriorityInput Tokens (per 1M)Cache Read Tokens (per 1M)Output Tokens (per 1M)
Realtime1$0.15$0.04$0.60
Async$0.11$0.03$0.45
Batch (24h)$0.08$0.02$0.30

Playground

Open this model in the Playground.

Footnotes

  1. Realtime availability is limited. Doubleword is primarily a batch API.