LLM Pricing Calculator
Compare API pricing across OpenAI, Anthropic, Google, Mistral, and more. Estimate self-hosting costs and find your break-even point.
Last updated: June 13, 2026
1.0M tokens ≈ 750K words ≈ 1,500 pages
Compare your own model pricing
| Provider / Model | Tier | Monthly Cost | $/1M Tokens | Context | Capabilities | |
|---|---|---|---|---|---|---|
Cohere3 models | from $0.07 | from $0.07 | ||||
Google10 models | from $0.14 | from $0.14 | ||||
OpenAI22 models | from $0.16 | from $0.16 | ||||
Mistral9 models | from $0.16 | from $0.16 | ||||
DeepSeek4 models | from $0.18 | from $0.18 | ||||
xAI6 models | from $0.29 | from $0.29 | ||||
Anthropic13 models | from $0.55 | from $0.55 | ||||
API vs Self-Hosting: Cost Comparison
Compare the total cost of using an API provider versus hosting an open-source model yourself. Select models from the API pricing table to compare, or choose below.
Self-Hosting Configuration
1.0M tokens/mo
From your usage settings above
GPU Calculation
Precision
FP16
Throughput per GPU
330 tok/s
Capacity per GPU
855M/mo
GPUs Needed
1(1.0M ÷ 855M = 0.00)
VRAM per GPU
14 GB / 80 GB
Cheapest: Vast.ai A100 SXM @ $0.22/hr
$158.40/mo (GPU only)
Throughput: ~330 tok/s aggregate (vLLM, BS8). Varies with config.
Self-Hosting Overhead
Infrastructure setup, engineering time (adjust to your estimate)
Prompt management, monitoring, maintenance, on-call
Self-Hosted GPU
$158.40
+ Ops Overhead
$0.0000
Total Self-Hosted
$158.40/mo
Cheapest API
$0.07/mo
Command R7B
Cumulative Cost Over 12 Months
Showing top 5 cheapest API models vs self-hosted 7B.
Key Considerations
Latency
Data Privacy
Customization
Reliability
Disclaimer
Prices are approximate and based on publicly available pricing pages as of 2026-06-13. Actual costs may vary based on volume discounts, reserved pricing, and provider-specific terms. Always verify current pricing on provider websites before making purchasing decisions.
Frequently Asked Questions
How much does it cost to use the OpenAI API?
OpenAI API pricing varies by model. As of the calculator's last-verified pricing, GPT-5.5 costs $5 per million input tokens and $30 per million output tokens, GPT-5 costs $1.25/$10, and GPT-5 Nano costs $0.05/$0.40. The calculator applies the current per-token prices to your own monthly volumes, so you see a cost for each model rather than a generic estimate.
Is it cheaper to self-host an LLM or use an API?
At low to moderate volumes, APIs are typically cheaper because you only pay for the tokens you use. Self-hosting becomes cost-effective when sustained high throughput spreads fixed GPU costs over a large volume of tokens. The break-even point depends on your model size, utilization, and GPU rates; the calculator computes it from your own numbers.
What is the cheapest LLM API in 2026?
Among the models the calculator tracks, the lowest-priced as of its last-verified pricing include Cohere Command R7B ($0.0375 per million input tokens / $0.15 per million output), GPT-5 Nano ($0.05/$0.40), Llama 3.1 8B on Groq ($0.06/$0.06), and Mistral Small 4 ($0.10/$0.30). The best value depends on your quality requirements and use case; the calculator lets you compare within a model tier.
How do I calculate LLM API costs for my project?
Estimate your monthly token usage (1 token ≈ 0.75 words), split between input and output, then multiply by each per-million-token price. For example, 5M tokens/month at a 60/40 input/output split on GPT-5.4 ($2.50 input / $15 output per million tokens, as of the calculator's last-verified pricing): 3M input × $2.50/1M + 2M output × $15/1M = $7.50 + $30 = $37.50/month.