DoublewordDoubleword

Model Name

thinkingmachines/Inkling-NVFP4

Inkling

  • Type: Generation
  • Capabilities: vision, reasoning
  • Cache read: $0.20 per 1M input tokens (0.17× standard input price). See prompt caching.

Overview

Inkling is the first model released by Thinking Machines - a large Mixture of Experts model with 975B parameters, 41B active, trained on over 45T tokens of text, image, and audio data. It supports up to a 1M context length and and offers strong performance across the board for agentic and reasoning heavy tasks.

Pricing

PriorityInput Tokens (per 1M)Cache Read Tokens (per 1M)Output Tokens (per 1M)
Realtime1$1.20$0.20$4.00
Async$0.90$0.15$3.00
Batch (24h)$0.60$0.10$2.00

Playground

Open this model in the Playground.

Footnotes

  1. Realtime availability is limited. Doubleword is primarily a batch API.