Model Name
thinkingmachines/Inkling-NVFP4Inkling
- Type: Generation
- Capabilities:
vision,reasoning - Cache read: $0.20 per 1M input tokens (0.17× standard input price). See prompt caching.
Overview
Inkling is the first model released by Thinking Machines - a large Mixture of Experts model with 975B parameters, 41B active, trained on over 45T tokens of text, image, and audio data. It supports up to a 1M context length and and offers strong performance across the board for agentic and reasoning heavy tasks.
Pricing
| Priority | Input Tokens (per 1M) | Cache Read Tokens (per 1M) | Output Tokens (per 1M) |
|---|---|---|---|
| Realtime1 | $1.20 | $0.20 | $4.00 |
| Async | $0.90 | $0.15 | $3.00 |
| Batch (24h) | $0.60 | $0.10 | $2.00 |
Playground
Open this model in the Playground.
Footnotes
-
Realtime availability is limited. Doubleword is primarily a batch API. ↩