DoublewordDoubleword

Model Name

nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

Nemotron 3 Super 120B A12B

  • Type: Generation
  • Capabilities: reasoning
  • Cache read: $0.02 per 1M input tokens (0.3× standard input price). See prompt caching.

Overview

NVIDIA Nemotron 3 Super 120B A12B NVFP4 is an open hybrid Mamba-Transformer LatentMoE model with 120 billion total parameters and 12 billion active parameters, built for agentic reasoning workloads such as coding, planning, tool use, and long-context tasks. It sits in the same capability tier as Qwen3.5-122B non-reasoning and ahead of GPT-OSS-120B, while also delivering higher throughput.


In line with NVIDIA's guidance, we use temperature=1.0 and top_p=0.95 across all tasks and serving backends, including reasoning, tool calling, and general chat.

Reasoning efforts

  • Supported: none, minimal, low, medium, high, xhigh, max

See the reasoning effort guide for request examples.

Pricing

PriorityInput Tokens (per 1M)Cache Read Tokens (per 1M)Output Tokens (per 1M)
Realtime1$0.08$0.02$0.45
Async$0.06$0.02$0.34
Batch (24h)$0.04$0.01$0.23

Playground

Open this model in the Playground.

Footnotes

  1. Realtime availability is limited. Doubleword is primarily a batch API.