DoublewordDoubleword
Get started

Model Name

Qwen/Qwen3.8-27B-FP8

Qwen3.8 27B

  • Type: Generation
  • Capabilities: reasoning, vision
  • Cache read: $0.04 per 1M input tokens (0.1× standard input price). See prompt caching.

Overview

Qwen3.8-27B is a compact multimodal reasoning model from Alibaba’s Qwen family, designed for general-purpose reasoning, coding, tool use, and vision workloads. Its 262K-token context window supports long documents and extended agentic tasks, while the FP8 deployment offers efficient serving.


Thinking Mode:

This model reasons step-by-step before responding by default.

Set reasoning_effort to none to disable thinking. Other effort levels enable thinking but do not select graduated reasoning budgets.

Reasoning efforts

  • Supported: none, minimal, low, medium, high, xhigh, max

See the reasoning effort guide for request examples.

Pricing

PriorityInput Tokens (per 1M)Cache Read Tokens (per 1M)Output Tokens (per 1M)
Realtime1$0.45$0.04$3.00
Async$0.35$0.04$2.25
Batch (24h)$0.25$0.02$1.50

Playground

Open this model in the Playground.

Footnotes

  1. Realtime availability is limited. Doubleword is primarily a batch API. ↩