Skip to content
OpenSmartRoute
All vendors
Meta

Vendor

Meta models

8 models in the reference catalogue, vendor list prices per million tokens. 1 is routable on this deployment today.

Models

8

1 routable here

Cheapest (blended)

$0.06

Llama 3.1 8B Instruct

Median list price

$0.15

3:1 input to output per 1M tokens

Strongest

-

no published benchmarks

Every Meta model

Newest first. Prices are the vendor's published rates per million tokens; click a model for the full specification.

ModelReleasedContextInput / 1MOutput / 1MIntelligenceCapabilities
Llama Guard 4 12Bmeta-llama/llama-guard-4-12bApr 30, 2025164K$0.18$0.18-
open weights
Llama 4 Maverickmeta-llama/llama-4-maverickApr 5, 20251.0M$0.20$0.70-
tools open weights
Llama 4 Scoutmeta-llama/llama-4-scoutApr 5, 20251.3M$0.10$0.30-
tools open weights
Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instructDec 6, 2024131K$0.10$0.32-
tools open weightsroutable
Llama 3.2 1B Instructmeta-llama/llama-3.2-1b-instructSep 25, 202460K$0.03$0.20-
open weights
Llama 3.2 3B Instructmeta-llama/llama-3.2-3b-instructSep 25, 2024131K$0.05$0.33-
open weights
Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instructJul 23, 2024131K$0.40$0.40-
tools open weights
Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instructJul 23, 2024131K$0.05$0.08-
tools open weights

Frequently asked

What is the cheapest Meta model?
Llama 3.1 8B Instruct at $0.05 per 1M input tokens and $0.08 per 1M output tokens (vendor list price).
Which Meta model has the largest context window?
Llama 4 Scout accepts 1.3M tokens of context.
How do I route to Meta models with OpenSmartRoute?
Declare a target in targets.yaml with metadata.model set to the catalogue id (for example meta-llama/llama-guard-4-12b); the router scores it against every other target on cost, quality, latency and your policies for each request.

Route Meta with everything else

OpenSmartRoute picks the cheapest model that meets your quality bar per request, so a Meta flagship handles hard prompts while small models take the rest. Price a workload or try a routing decision live.