NVIDIA Models
NVIDIA's hosted inference for its Nemotron family of open models. Retains API data and may use it to train and improve models. Requesty routes to 7 NVIDIA models with context windows up to 1.0M tokens. One API key and an OpenAI-compatible SDK. Provider rates plus 5% on pay as you go, or 0% with your own keys.
Flagship model
nemotron-3-ultra-550b-a55bIntelligence Index
29.3
Coding Index
49.3
GPQA Diamond
86.7%
Terminal-Bench Hard
36.4%
All NVIDIA models
| Model | Context | Max Output | Input/1M | Output/1M | Capabilities | Coding |
|---|---|---|---|---|---|---|
| 1.0M | 66K | Free | Free | π§ π§ | N/A | |
| 131K | 20K | Free | Free | ππ§ π§ | N/A | |
| 131K | 8K | Free | Free | ππ§ | N/A | |
| 131K | 20K | Free | Free | ππ§ π§ | N/A | |
| 1.0M | 66K | Free | Free | π§ π§ | 49 | |
| 1.0M | 66K | Free | Free | π§ π§ | 38 | |
| 262K | β | Free | Free | π§ π§ | 14 |
About NVIDIA on Requesty
How many NVIDIA models are available through Requesty?
Requesty routes to 7 NVIDIA models including regional variants, with pricing synced in real time to the upstream provider.
What is the cheapest NVIDIA model?
NVIDIA has free tiers available: look for the models marked "Free" in the pricing column.
Does Requesty add markup on NVIDIA pricing?
No. Requesty passes through exactly what NVIDIA charges. You pay the same per-token rates as going direct, plus you get smart routing, caching, analytics, and one unified API for 600+ models.
Is my data used to train NVIDIA models?
NVIDIA's default terms may include data use for training. Check their privacy policy and Requesty's enterprise options for opt-out controls.
Where are NVIDIA models hosted?
NVIDIA models are hosted in πΊπΈ US. Some models are available in additional regions through AWS Bedrock, Azure, or Google Vertex AI: filter by region on the NVIDIA rows in the models explorer.
