Vendor
Moonshot AI models
8 models in the reference catalogue, vendor list prices per million tokens. 0 are routable on this deployment today.
Models
8
0 routable here
Median list price
$1.21
3:1 input to output per 1M tokens
Every Moonshot AI model
Newest first. Prices are the vendor's published rates per million tokens; click a model for the full specification.
| Model | Released | Context | Input / 1M | Output / 1M | Intelligence | Capabilities |
|---|---|---|---|---|---|---|
| Kimi K3moonshotai/kimi-k3 | Jul 16, 2026 | 1.0M | $3.00 | $15.0 | 50 | reasoning tools open weights |
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | Jun 12, 2026 | 262K | $0.66 | $3.40 | - | reasoning tools open weights |
| MoonshotAI Kimi Latest~moonshotai/kimi-latest | Apr 27, 2026 | 1.0M | $2.55 | $12.8 | - | reasoning tools |
| Kimi K2.6moonshotai/kimi-k2.6 | Apr 20, 2026 | 262K | $0.95 | $4.00 | - | reasoning tools open weights |
| Kimi K2.5moonshotai/kimi-k2.5 | Jan 27, 2026 | 262K | $0.45 | $2.25 | - | reasoning tools open weights |
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | Nov 6, 2025 | 262K | $0.60 | $2.50 | - | reasoning tools open weights |
| Kimi K2 0905moonshotai/kimi-k2-0905 | Sep 4, 2025 | 262K | $0.60 | $2.50 | - | tools open weights |
| Kimi K2 0711moonshotai/kimi-k2 | Jul 11, 2025 | 131K | $0.57 | $2.30 | - | tools open weights |
Frequently asked
- What is the cheapest Moonshot AI model?
- Kimi K2.5 at $0.45 per 1M input tokens and $2.25 per 1M output tokens (vendor list price).
- Which Moonshot AI model has the largest context window?
- Kimi K3 accepts 1.0M tokens of context.
- How do I route to Moonshot AI models with OpenSmartRoute?
- Declare a target in targets.yaml with metadata.model set to the catalogue id (for example moonshotai/kimi-k3); the router scores it against every other target on cost, quality, latency and your policies for each request.
Route Moonshot AI with everything else
OpenSmartRoute picks the cheapest model that meets your quality bar per request, so a Moonshot AI flagship handles hard prompts while small models take the rest. Price a workload or try a routing decision live.