Model Name
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4Nemotron 3 Super 120B A12B
- Type: Generation
- Capabilities:
reasoning - Cache read: $0.02 per 1M input tokens (0.3× standard input price). See prompt caching.
Overview
NVIDIA Nemotron 3 Super 120B A12B NVFP4 is an open hybrid Mamba-Transformer LatentMoE model with 120 billion total parameters and 12 billion active parameters, built for agentic reasoning workloads such as coding, planning, tool use, and long-context tasks. It sits in the same capability tier as Qwen3.5-122B non-reasoning and ahead of GPT-OSS-120B, while also delivering higher throughput.
In line with NVIDIA's guidance, we use temperature=1.0 and top_p=0.95 across all tasks and serving backends, including reasoning, tool calling, and general chat.
Reasoning efforts
- Supported:
none,minimal,low,medium,high,xhigh,max
See the reasoning effort guide for request examples.
Pricing
| Priority | Input Tokens (per 1M) | Cache Read Tokens (per 1M) | Output Tokens (per 1M) |
|---|---|---|---|
| Realtime1 | $0.08 | $0.02 | $0.45 |
| Async | $0.06 | $0.02 | $0.34 |
| Batch (24h) | $0.04 | $0.01 | $0.23 |
Playground
Open this model in the Playground.
Footnotes
-
Realtime availability is limited. Doubleword is primarily a batch API. ↩