Mistral API Pricing 2026: The Overlooked Budget Powerhouse

Here is a number that should make you reconsider your LLM bill: Mistral Large 3, the company's open-weight flagship, costs $0.50 per million input tokens and $1.50 per million output. That is roughly 80% cheaper on input than OpenAI's GPT-5.6 ($2.50) and 90% cheaper than Anthropic's Claude Opus 4.8 ($5). The output gap is wider still.
Mistral is a French AI company that ships competitive open and premier models at aggressive price points, and most engineering teams we work with have never seriously evaluated them. This guide breaks down Mistral API pricing for every model as of July 2026, across text, reasoning, code, vision, voice, and OCR. For each, you get what it costs and where it makes sense against the big three.
All token prices are per million tokens. Run your own numbers in our LLM Pricing Calculator.
The Full Mistral API Pricing Table (July 2026)
Text and reasoning models, priced per million tokens:
| Model | Input (/1M) | Output (/1M) | Context | Best for |
| Mistral Large 3 | $0.50 | $1.50 | 256K | Open-weight flagship, general + vision |
| Mistral Medium 3.5 | $1.50 | $7.50 | 256K | Enterprise-grade quality & deployment |
| Mistral Small 4 | $0.15 | $0.60 | 256K | High-volume, multimodal, Apache 2.0 |
| Magistral Medium | $2.00 | $5.00 | 128K | Deep, transparent reasoning |
| Magistral Small | $0.50 | $1.50 | 128K | Budget reasoning |
| Codestral | $0.30 | $0.90 | 256K | Low-latency code completion |
| Devstral 2 | $0.40 | $2.00 | 256K | Agentic coding |
| Devstral Small 2 | $0.10 | $0.30 | 256K | Lightweight coding agents |
| Ministral 3 · 3B | $0.10 | $0.10 | 128K | Edge / on-device |
| Ministral 3 · 8B | $0.15 | $0.15 | 256K | Edge / on-device |
| Ministral 3 · 14B | $0.20 | $0.20 | 256K | Edge / on-device |
| Mistral NeMo | $0.15 | $0.15 | 128K | Legacy open lightweight |
Specialized services (documents, voice, embeddings, moderation):
| Service | Price |
| Mistral OCR 4 — OCR | $4 / 1,000 pages |
| Mistral OCR 4 — Document AI | $5 / 1,000 pages |
| Voxtral TTS (text-to-speech) | $0.016 / 1,000 characters |
| Voxtral Mini Transcribe 2 | $0.003 / minute of audio |
| Voxtral Small (audio + text) | $0.004/min audio · $0.10 in / $0.40 out per 1M |
| Codestral Embed | $0.15 / 1M tokens |
| Mistral Embed | $0.10 / 1M tokens |
| Mistral Moderation | $0.10 / 1M tokens |
Agent and tool calls, billed per use on top of model tokens:
| Tool | Price |
| Web search | $30 / 1,000 calls |
| Code execution | $30 / 1,000 calls |
| Image generation | $100 / 1,000 images |
| Premium news | $50 / 1,000 calls |
| Document libraries | $3/1K pages OCR + $1/1M indexing + $0.01/call |
Vision is now built into Large 3, Medium 3.5, and Small 4, so the standalone Pixtral Large model has been retired. Leanstral, an open Lean 4 code agent, is currently free.
Why Mistral Deserves Your Attention
The pricing tells the story immediately. Here is how the flagships compare on a representative 5M-input / 2M-output workload:
| Provider | Flagship | Input | Output | Total (5M in / 2M out) |
| Mistral | Large 3 | $0.50 | $1.50 | $5.50 |
| Gemini 3.1 Pro | $2.00 | $12.00 | $34.00 | |
| OpenAI | GPT-5.6 | $2.50 | $15.00 | $42.50 |
| Anthropic | Claude Opus 4.8 | $5.00 | $25.00 | $75.00 |
Mistral Large 3 runs that workload for $5.50 — versus $34 on Gemini 3.1 Pro, $42.50 on GPT-5.6, and $75 on Claude Opus 4.8. That is roughly 8x cheaper than OpenAI's flagship and 14x cheaper than Anthropic's, for the same tokens.

Bar chart: monthly cost of a 5M-input / 2M-output workload — Mistral Large 3 $5.50 vs Google Gemini 3.1 Pro $34, OpenAI GPT-5.6 $42.50, Anthropic Claude Opus 4.8 $75
Monthly cost for the same 5M-input / 2M-output workload across flagship models, from official provider pricing (July 2026).
Raw price only matters if quality holds. Mistral Large 3 benchmarks competitively with previous-generation flagships and, as an open-weight model, can also be self-hosted. It may not match GPT-5.6 or Claude Opus 4.8 on the hardest tasks, but for a large share of production workloads the quality gap is smaller than the price gap.
Model Deep Dives
Mistral Large 3: The Value Flagship
At $0.50/$1.50, Large 3 offers one of the best price-to-quality ratios on the market. It is open-weight, multimodal (text + vision), supports function calling, and can be run through the API or self-hosted.
The output price is the standout: $1.50 per million tokens, against $15 for GPT-5.6, $25 for Claude Opus 4.8, and $10 for Claude Sonnet 5. If your workload is output-heavy (content, code, or report generation), Large 3 cuts output cost by 85–94%.
Best for:
- General-purpose production workloads where you are currently overpaying a premium provider
- Output-heavy pipelines (content, code, reports) where the $1.50 output price compounds
- Multilingual applications — Mistral models are strong across European languages
- Teams that want the option to self-host an open-weight flagship
Mistral Medium 3.5: The Enterprise Tier
Counterintuitively, Medium 3.5 ($1.50/$7.50) costs more than the Large 3 flagship. It targets enterprise deployments, tuned for quality and simplified private hosting rather than lowest cost. For most API users optimizing spend, Large 3 is the better pick; reach for Medium 3.5 when its specific quality profile or deployment story matters to you.
Mistral Small 4: The Volume Workhorse
At $0.15/$0.60, Small 4 competes with Gemini 2.0 Flash ($0.15/$0.60), Gemini 2.5 Flash-Lite ($0.10/$0.40), and GPT-5.4 Mini ($0.75/$4.50), undercutting the OpenAI option heavily on output. It is Apache-2.0 licensed, multimodal, and supports function calling, which is rare at this price.
Ministral 3: Frontier AI at the Edge
New in the current lineup, the Ministral 3 family (3B, 8B, 14B) is built for edge and on-device use, with flat pricing where input and output cost the same: $0.10 (3B), $0.15 (8B), and $0.20 (14B) per million tokens. For latency- or privacy-sensitive workloads that run close to the user, these are the cheapest competent options Mistral offers.
Magistral: Mistral's Reasoning Models
The Magistral line uses extended chain-of-thought for complex problems. Magistral Medium ($2.00/$5.00) and Magistral Small ($0.50/$1.50) deliver transparent, multilingual reasoning. Since OpenAI folded reasoning into the GPT-5 family (GPT-5.4 at $2.50/$15), Magistral Medium's $5 output undercuts it by two-thirds; Magistral Small matches Large 3's price while optimizing for reasoning.
Codestral & Devstral: Purpose-Built for Code
Codestral ($0.30/$0.90, 256K context) is a low-latency model for completion, fill-in-the-middle, and generation. Devstral 2 ($0.40/$2.00) and Devstral Small 2 ($0.10/$0.30) are agentic coding models for autonomous software engineering, and Leanstral is a free open code agent for Lean 4. Against GPT-5.6 ($2.50/$15) or Claude Sonnet 5 ($2/$10) for code, Codestral is 6–15x cheaper on the same tokens.
Vision, Voice, and Documents
Vision is now built into Large 3, Medium 3.5, and Small 4, so the dedicated Pixtral Large model has been retired and you no longer pay a premium for multimodal input. For speech, Voxtral covers text-to-speech ($0.016 / 1,000 characters) and transcription (from $0.003 / minute). For documents, Mistral OCR 4 runs at $4 / 1,000 pages ($5 with Document AI extraction), priced per page rather than per token.
Embeddings & Moderation
Mistral Embed ($0.10 / 1M) and the code-specialized Codestral Embed ($0.15 / 1M) cover retrieval and RAG, and Mistral Moderation ($0.10 / 1M) provides content classification, all at commodity prices.
Real Cost Calculations
Use Case 1 — AI Email Drafting
20,000 drafts per month, roughly 800 input + 400 output tokens each: 16M input, 8M output.
| Model | Monthly Cost | Per-Email |
| Mistral Small 4 | $7.20 | $0.00036 |
| Mistral Large 3 | $20.00 | $0.001 |
| GPT-5.4 Mini (OpenAI) | $48.00 | $0.0024 |
| Claude Sonnet 5 | $112.00 | $0.0056 |
Small 4 undercuts everyone at $7.20/month; Large 3 at $20 still beats GPT-5.4 Mini by more than half. Claude Sonnet 5 at $112 is about 16x Small 4 for this shape of work.
Use Case 2 — Code Review Pipeline
500 PRs per month, roughly 6,000 input + 1,500 output tokens each: 3M input, 0.75M output.
| Model | Monthly Cost |
| Codestral | $1.58 |
| Mistral Large 3 | $2.63 |
| Claude Sonnet 5 | $13.50 |
| GPT-5.6 (OpenAI) | $18.75 |
Codestral at $1.58/month for 500 reviews is remarkably cheap: an 8–12x cost advantage over Sonnet 5 and GPT-5.6. Even at 80% of their quality it is worth a serious evaluation.
Use Case 3 — Reasoning-Heavy Analysis
1,000 tasks per month, roughly 5,000 input + 3,000 output tokens each: 5M input, 3M output.
| Model | Monthly Cost |
| Magistral Small | $7.00 |
| Magistral Medium | $25.00 |
| Claude Sonnet 5 (extended thinking) | $40.00 |
| Gemini 3.1 Pro | $46.00 |
| GPT-5.4 (OpenAI) | $57.50 |
Magistral Small delivers reasoning at $7/month, roughly one-eighth of GPT-5.4. Magistral Medium at $25 still lands under every premium option while giving a larger reasoning budget.
Where Mistral Falls Short
Transparency matters, so here is where Mistral is not the right pick:
Context window: Mistral's flagships now reach 256K tokens (Magistral and the smallest edge model stay at 128K). That is a real jump from the 128K of a year ago, but still short of GPT-5.6's 400K or the 1M windows on Gemini 3.1 Pro and Claude Sonnet 5 / Opus 4.8.
Ecosystem and tooling: OpenAI, Google, and Anthropic have larger ecosystems, with more SDKs, integrations, and community. Switching cost is real if you depend on them.
Peak instruction-following: Claude Sonnet 5 and Opus 4.8 remain the reference for long, nuanced, ambiguous instructions. If precision on complex prompts is critical, Mistral may need more prompt engineering to match.
Enterprise features: verify Mistral's current SOC 2, DPA, and SLA coverage against your requirements; the incumbents have more mature programs.
Benchmarks vs reality: competitive benchmark scores do not guarantee production performance. Always run your own evaluations.
The Migration Path
If you are considering a move from OpenAI or Anthropic to Mistral:
- Pick your highest-volume, lowest-complexity workload. Classification, extraction, and simple generation are the best candidates.
- Run a parallel evaluation: send the same 200–500 requests to both your current model and the Mistral equivalent, and score on your real quality criteria.
- Calculate the real savings: token cost plus latency, error rates, and any prompt engineering needed to adapt.
- Migrate the easy wins; keep the harder tasks on your current provider.
- Re-evaluate quarterly, since Mistral ships new models frequently.
Most teams we work with end up multi-provider: Mistral for volume, Claude or GPT-5.6 for the tasks that need peak quality. Blended cost is typically 40–60% lower than running a single premium provider for everything — see our full OpenAI vs Gemini vs Claude pricing comparison for the head-to-head.
The Bottom Line
Mistral is not a compromise. It is a legitimate alternative that happens to cost 80–95% less than the incumbents for many workloads. The question is not “is Mistral good enough” — it is “have you actually tested it on your workload?”
If you have not, you may be leaving real money on the table.
Prices verified against Mistral's official API pricing page (mistral.ai/pricing) as of July 2026. Compare Mistral against every provider in our LLM Pricing Calculator.
Sources
Pricing verified against official provider pricing pages, retrieved 2026-07-17:
Need help building AI into your product?
We design, build, and integrate production AI systems. Talk directly with the engineers who'll build your solution.
Get in touchWritten by
Aniket Kulkarni
Aniket Kulkarni is the founder of Curlscape, an AI consulting firm that helps companies build and ship production AI systems. With experience spanning voice agents, LLM evaluation harnesses, and bespoke AI solutions, he works at the intersection of engineering and applied machine learning. He writes about practical AI implementation, model selection, and the tools shaping the AI ecosystem.
Frequently Asked Questions
Is Mistral as good as ChatGPT or Claude?▼
Mistral Large 3 benchmarks competitively with recent flagships at a fraction of the cost — $0.50/$1.50 versus GPT-5.6 at $2.50/$15 or Claude Opus 4.8 at $5/$25. It may not match the very best models on the hardest tasks, but for classification, extraction, and content generation the quality gap is usually smaller than the 8–14x price gap. Always benchmark on your own use case.
What is the cheapest Mistral model?▼
For general use, Mistral Small 4 at $0.15/$0.60 per million tokens — multimodal, function-calling, Apache 2.0. For the absolute lowest cost, the edge-focused Ministral 3 3B is $0.10 flat for both input and output, and Devstral Small 2 is $0.10/$0.30 for coding agents.
Should I use Codestral or a general model for code?▼
Codestral ($0.30/$0.90, 256K context) is purpose-built for code completion and generation, and is cheaper and often better than a general model for pure code tasks. For autonomous, multi-step coding agents, look at Devstral 2 ($0.40/$2.00). Use a general model like Large 3 when you need mixed code-plus-explanation output.
Does Mistral have reasoning models?▼
Yes — the Magistral line. Magistral Medium ($2.00/$5.00) and Magistral Small ($0.50/$1.50) use extended chain-of-thought for complex, multilingual reasoning, at a fraction of the cost of premium reasoning from the larger providers.
Can I use Mistral as a drop-in replacement for OpenAI?▼
Largely yes. Mistral's API follows the OpenAI-compatible chat-completions format, so migration is mostly swapping the endpoint and key. You will likely need to adjust prompts, since models respond differently to the same instructions, but the engineering effort is typically small.
Continue Reading

Google Gemini API Pricing Guide 2026: Flash, Pro, and Vertex AI
Current Google Gemini API pricing for 2026: Gemini 3 generation (3.1 Pro, 3.5 Flash, 3.1 Flash-Lite), what changed since 2.5, image generation with Nano Banana, and how Vertex AI costs compare.

Anthropic Claude API Pricing Guide 2026: Opus, Sonnet, and Haiku Compared
Complete Anthropic Claude API pricing for March 2026. Compare Opus, Sonnet 4.6, and Haiku 4.5 with batch discounts, prompt caching savings, rate limits, and real-world cost breakdowns.

Fine-tuning open models in the real world: Unsloth, Axolotl, and the case for Docker
Production lessons from fine-tuning open models and why Curlscape uses Docker to ensure GPU training environments are reproducible and reliable.