Requesty
Thinking Machines logo

Thinking Machines Models

US AI lab serving its own open-weight Inkling models (text, image and audio input). Inputs are retained as needed to provide the service and are not used to train models. Requesty routes to 2 Thinking Machines models starting at $1.87 per 1M input tokens with context windows up to 262K tokens. One API key and an OpenAI-compatible SDK. Provider rates plus 5% on pay as you go, or 0% with your own keys.

Thinking Machines homepage →
Intelligence Index
25.0
Coding Index
52.1
GPQA Diamond
87.2%

All Thinking Machines models

ModelContextMax OutputInput/1MOutput/1MCapabilitiesCoding
66K33K$1.87$4.68
👁🧠🔧⚡
52
262K33K$3.74$1.87$9.36$4.68
👁🧠🔧⚡
N/A

About Thinking Machines on Requesty

How many Thinking Machines models are available through Requesty?
Requesty routes to 2 Thinking Machines models including regional variants, with pricing synced in real time to the upstream provider.
What is the cheapest Thinking Machines model?
The cheapest Thinking Machines model starts at $1.87 per million input tokens. See the pricing column in the table below for full per-model rates.
Does Requesty add markup on Thinking Machines pricing?
No. Requesty passes through exactly what Thinking Machines charges. You pay the same per-token rates as going direct, plus you get smart routing, caching, analytics, and one unified API for 600+ models.
Is my data used to train Thinking Machines models?
Thinking Machines's terms state that API data is not used for training. See their privacy policy for the authoritative statement.
Where are Thinking Machines models hosted?
Thinking Machines models are hosted in 🇺🇸 US. Some models are available in additional regions through AWS Bedrock, Azure, or Google Vertex AI: filter by region on the Thinking Machines rows in the models explorer.