Model Name
Qwen/Qwen3-VL-30B-A3B-Instruct-FP8Qwen3 VL 30B A3B Instruct
- Type: Generation
- Capabilities:
vision - Cache read: $0.04 per 1M input tokens (0.25× standard input price). See prompt caching.
Overview
Meet Qwen3-VL-30B, the smaller model of the Qwen3-VL family, delivering performance similar to GPT-4.1-mini and Claude Sonnet 4. This highly capable mid-size model is suited for tasks that are constrained or require high token volumes. Excels at reasoning, coding, and structured output generation.
Best for:
- Production workloads requiring strong performance without frontier model costs
- Complex reasoning tasks
- Code generation
Pricing
| Priority | Input Tokens (per 1M) | Cache Read Tokens (per 1M) | Output Tokens (per 1M) |
|---|---|---|---|
| Realtime1 | $0.15 | $0.04 | $0.60 |
| Async | $0.11 | $0.03 | $0.45 |
| Batch (24h) | $0.08 | $0.02 | $0.30 |
Playground
Open this model in the Playground.
Footnotes
-
Realtime availability is limited. Doubleword is primarily a batch API. ↩