AI APIs · Price comparison
AI API pricing comparison
Compare public list prices across 11 current AI models from five providers. Every number below is priced per 1 million tokens and links back to a source. Use the calculator for your own input/output mix — the cheapest input rate is not always the cheapest workload.
What is the cheapest AI API here?
Cheapest input tokens
Qwen3.5 Flash
$0.065 per 1M input tokens
Cheapest output tokens
Qwen3.5 Flash
$0.26 per 1M output tokens
This is a rate-card answer, not a quality ranking. A model that needs more output tokens, retries, or human correction can cost more per completed task even when its listed token price is lower.
LLM API price comparison per 1M tokens
| Model | Provider | Input / 1M | Output / 1M | Cached input | Context |
|---|---|---|---|---|---|
| Qwen3.5 Flash | Alibaba | $0.065 | $0.26 | — | 1M |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | — | Not stated | |
| Qwen3.5 Plus | Alibaba | $0.30 | $1.80 | — | 1M |
| Qwen3.7 Plus | Alibaba | $0.32 | $1.28 | — | 1M |
| GPT-5.6 Luna | OpenAI | $1.00 | $6.00 | — | 1M |
| Qwen3.7 Max | Alibaba | $1.48 | $4.43 | — | 1M |
| Gemini 3.6 Flash | $1.50 | $7.50 | — | Not stated | |
| GPT-5.6 Terra | OpenAI | $2.50 | $15.00 | — | 1M |
| Kimi K3 | Moonshot AI | $3.00 | $15.00 | $0.30 | 1M |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | — | 1M |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | — | Not stated |
Sorted by input price. Prices exclude batch discounts, provider promotions, enterprise agreements and tool-call charges. “Not stated” means the cited source did not publish a context window; it is not an estimate.
AI token cost calculator
Enter the same workload for every model. The calculator combines input and output cost and marks the lowest list-price total for that token mix.
Cost calculator
Enter your monthly token volume to see what each tier would cost.
| Tier | Input cost | Output cost | Monthly total |
|---|---|---|---|
| Qwen3.5 Flash | $0.065 | $0.052 | $0.117cheapest |
| Gemini 3.5 Flash-Lite | $0.30 | $0.50 | $0.80 |
| Qwen3.5 Plus | $0.30 | $0.36 | $0.66 |
| Qwen3.7 Plus | $0.32 | $0.256 | $0.576 |
| Luna | $1.00 | $1.20 | $2.20 |
| Qwen3.7 Max | $1.48 | $0.885 | $2.36 |
| Gemini 3.6 Flash | $1.50 | $1.50 | $3.00 |
| Terra | $2.50 | $3.00 | $5.50 |
| Kimi K3 | $3.00 | $3.00 | $6.00 |
| Sol | $5.00 | $6.00 | $11.00 |
| Opus 5 | $5.00 | $5.00 | $10.00 |
Estimates only — list prices, before any caching discount, batch pricing, or enterprise agreement.
How to compare AI API pricing without fooling yourself
Separate input and output
Output is often several times more expensive. A chat app and a document classifier can rank models differently even at the same total token volume.
Measure tokens per accepted result
Reasoning verbosity, retries and rejected answers change the real bill. Benchmark the task you actually run, not only the vendor's per-token rate.
Recheck the rate card
Vendors change prices and promotions. Each row carries sourced data, and the page shows when this comparison was last verified.
Need detail on one family? Start with the GPT-5.6 price breakdown or the Kimi K3 cost analysis.
Sources
- OpenRouter — Qwen model catalogue — retrieved 2026-07-20
- Google — Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber (launch post) — retrieved 2026-07-22
- OpenRouter — GPT-5.6 Luna — retrieved 2026-07-17
- OpenRouter — GPT-5.6 Terra — retrieved 2026-07-17
- OpenRouter — Kimi K3 — retrieved 2026-07-21
- Artificial Analysis — Kimi K3 achieves #3 in the Intelligence Index — retrieved 2026-07-21
- Artificial Analysis — Kimi K3 model page (live index and price) — retrieved 2026-07-25
- Hugging Face — moonshotai organisation (checked: no Kimi K3 repository) — retrieved 2026-07-25
- Hugging Face — Kimi-K2.7-Code model card (Modified MIT License precedent) — retrieved 2026-07-25
- Tom's Hardware — Kimi K3 beats Claude Fable 5 in Frontend Code Arena — retrieved 2026-07-25
- The Decoder — Alibaba's Qwen takes on Kimi K3 with open-weight Qwen 3.8 — retrieved 2026-07-21
- Hugging Face community post — Kimi K3 architecture, MXFP4 quantization and weight-release notes — retrieved 2026-07-21
- OpenRouter — GPT-5.6 Sol — retrieved 2026-07-17
- Anthropic — Introducing Claude Opus 5 (launch post, rates and Fast Mode) — retrieved 2026-07-25
- Artificial Analysis — Opus 5: Fable 5 level intelligence at a lower cost per task — retrieved 2026-07-25
- Artificial Analysis — Claude Opus 5 (max) model page — retrieved 2026-07-25
- Implicator.ai — Opus 5 cut tokens 17% but cost more to benchmark than Opus 4.8 — retrieved 2026-07-25
Figures on this page last checked against these sources on 2026-07-25. Vendors change pricing and specs without notice — if a number here disagrees with the vendor's own page, trust the vendor.