Vendor
Meta models
8 models in the reference catalogue, vendor list prices per million tokens. 1 is routable on this deployment today.
Models
8
1 routable here
Median list price
$0.15
3:1 input to output per 1M tokens
Strongest
-
no published benchmarks
Every Meta model
Newest first. Prices are the vendor's published rates per million tokens; click a model for the full specification.
| Model | Released | Context | Input / 1M | Output / 1M | Intelligence | Capabilities |
|---|---|---|---|---|---|---|
| Llama Guard 4 12Bmeta-llama/llama-guard-4-12b | Apr 30, 2025 | 164K | $0.18 | $0.18 | - | open weights |
| Llama 4 Maverickmeta-llama/llama-4-maverick | Apr 5, 2025 | 1.0M | $0.20 | $0.70 | - | tools open weights |
| Llama 4 Scoutmeta-llama/llama-4-scout | Apr 5, 2025 | 1.3M | $0.10 | $0.30 | - | tools open weights |
| Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | Dec 6, 2024 | 131K | $0.10 | $0.32 | - | tools open weightsroutable |
| Llama 3.2 1B Instructmeta-llama/llama-3.2-1b-instruct | Sep 25, 2024 | 60K | $0.03 | $0.20 | - | open weights |
| Llama 3.2 3B Instructmeta-llama/llama-3.2-3b-instruct | Sep 25, 2024 | 131K | $0.05 | $0.33 | - | open weights |
| Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | Jul 23, 2024 | 131K | $0.40 | $0.40 | - | tools open weights |
| Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | Jul 23, 2024 | 131K | $0.05 | $0.08 | - | tools open weights |
Frequently asked
- What is the cheapest Meta model?
- Llama 3.1 8B Instruct at $0.05 per 1M input tokens and $0.08 per 1M output tokens (vendor list price).
- Which Meta model has the largest context window?
- Llama 4 Scout accepts 1.3M tokens of context.
- How do I route to Meta models with OpenSmartRoute?
- Declare a target in targets.yaml with metadata.model set to the catalogue id (for example meta-llama/llama-guard-4-12b); the router scores it against every other target on cost, quality, latency and your policies for each request.
Route Meta with everything else
OpenSmartRoute picks the cheapest model that meets your quality bar per request, so a Meta flagship handles hard prompts while small models take the rest. Price a workload or try a routing decision live.