Model Name
openai/gpt-oss-120bGPT OSS 120B
- Type: Generation
- Capabilities:
reasoning - Cache read: $0.11 per 1M input tokens (0.75× standard input price). See prompt caching.
Overview
Meet gpt-oss-20b — OpenAI’s larger open-weight MoE model, with 117B total parameters, 5.1B active parameters, and a 128k-token context window.
Reasoning efforts
- Supported:
low,medium,high
See the reasoning effort guide for request examples.
Pricing
| Priority | Input Tokens (per 1M) | Cache Read Tokens (per 1M) | Output Tokens (per 1M) |
|---|---|---|---|
| Realtime1 | $0.15 | $0.11 | $0.60 |
| Async | $0.11 | $0.08 | $0.45 |
| Batch (24h) | $0.08 | $0.06 | $0.30 |
Playground
Open this model in the Playground.
Footnotes
-
Realtime availability is limited. Doubleword is primarily a batch API. ↩