Vendor
NVIDIA models
10 models in the reference catalogue, vendor list prices per million tokens. 0 are routable on this deployment today.
Models
10
0 routable here
Median list price
$0.04
3:1 input to output per 1M tokens
Strongest
-
no published benchmarks
Every NVIDIA model
Newest first. Prices are the vendor's published rates per million tokens; click a model for the full specification.
| Model | Released | Context | Input / 1M | Output / 1M | Intelligence | Capabilities |
|---|---|---|---|---|---|---|
| Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightning | Aug 11, 2026 | 262K | $0.08 | $0.20 | - | reasoning tools open weights |
| Nemotron 3.5 Lightning (free)nvidia/nemotron-3.5-lightning:free | Aug 11, 2026 | 1M | free | free | - | reasoning tools open weights |
| Nemotron 3.5 Content Safetynvidia/nemotron-3.5-content-safety | Jun 4, 2026 | 131K | $0.20 | $0.20 | - | reasoning open weights |
| Nemotron 3.5 Content Safety (free)nvidia/nemotron-3.5-content-safety:free | Jun 4, 2026 | 128K | free | free | - | reasoning open weights |
| Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | Jun 4, 2026 | 262K | $0.63 | $3.13 | - | reasoning tools open weights |
| Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | Jun 4, 2026 | 1M | free | free | - | reasoning tools open weights |
| Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | Apr 28, 2026 | 256K | free | free | - | reasoning tools open weights |
| Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b | Mar 11, 2026 | 1M | $0.09 | $0.40 | - | reasoning tools open weights |
| Nemotron 3 Super (free)nvidia/nemotron-3-super-120b-a12b:free | Mar 11, 2026 | 262K | free | free | - | reasoning tools open weights |
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | Dec 14, 2025 | 262K | $0.05 | $0.20 | - | reasoning tools open weights |
Frequently asked
- What is the cheapest NVIDIA model?
- Nemotron 3.5 Lightning (free) at free per 1M input tokens and free per 1M output tokens (vendor list price).
- Which NVIDIA model has the largest context window?
- Nemotron 3.5 Lightning (free) accepts 1M tokens of context.
- How do I route to NVIDIA models with OpenSmartRoute?
- Declare a target in targets.yaml with metadata.model set to the catalogue id (for example nvidia/nemotron-3.5-lightning); the router scores it against every other target on cost, quality, latency and your policies for each request.
Route NVIDIA with everything else
OpenSmartRoute picks the cheapest model that meets your quality bar per request, so a NVIDIA flagship handles hard prompts while small models take the rest. Price a workload or try a routing decision live.