Curlscape logo

LLM Pricing Calculator

Compare API pricing across OpenAI, Anthropic, Google, Mistral, and more. Estimate self-hosting costs and find your break-even point.

Last updated: June 13, 2026

Pricing data last verified: 2026-06-1367 API Models14 GPU Options
1.0M tokens/mo
100K1M10M100MCustom

1.0M tokens ≈ 750K words ≈ 1,500 pages

Not sure? 70/30 input/output is a common starting point for chat workloads.
70% / 30%
More Input (lower cost)More Output (higher cost)
|

Compare your own model pricing

API vs Self-Hosting: Cost Comparison

Compare the total cost of using an API provider versus hosting an open-source model yourself. Select models from the API pricing table to compare, or choose below.

Self-Hosting Configuration

Token Demand

1.0M tokens/mo

From your usage settings above

GPU Calculation

Precision

FP16

Throughput per GPU

330 tok/s

Capacity per GPU

855M/mo

GPUs Needed

1(1.0M ÷ 855M = 0.00)

VRAM per GPU

14 GB / 80 GB

Cheapest: Vast.ai A100 SXM @ $0.22/hr

$158.40/mo (GPU only)

Throughput: ~330 tok/s aggregate (vLLM, BS8). Varies with config.

Self-Hosting Overhead

$

Infrastructure setup, engineering time (adjust to your estimate)

$/mo

Prompt management, monitoring, maintenance, on-call

Self-Hosted GPU

$158.40

+ Ops Overhead

$0.0000

Total Self-Hosted

$158.40/mo

Cheapest API

$0.07/mo

Command R7B

Cumulative Cost Over 12 Months

Showing top 5 cheapest API models vs self-hosted 7B.

Solid lines: API providersDashed green: Self-hosted (7B, 1 GPU)Self-hosted includes $5,000 setup cost

Key Considerations

Latency

APILow, managed by provider
SelfPredictable, no third-party network round-trip

Data Privacy

APIData sent to third party
SelfFull data sovereignty

Customization

APILimited to provider options
SelfFine-tuning, custom models

Reliability

APIProvider-managed uptime (SLAs on enterprise tiers)
SelfYou manage uptime

Disclaimer

Prices are approximate and based on publicly available pricing pages as of 2026-06-13. Actual costs may vary based on volume discounts, reserved pricing, and provider-specific terms. Always verify current pricing on provider websites before making purchasing decisions.

Frequently Asked Questions

How much does it cost to use the OpenAI API?

OpenAI API pricing varies by model. As of the calculator's last-verified pricing, GPT-5.5 costs $5 per million input tokens and $30 per million output tokens, GPT-5 costs $1.25/$10, and GPT-5 Nano costs $0.05/$0.40. The calculator applies the current per-token prices to your own monthly volumes, so you see a cost for each model rather than a generic estimate.

Is it cheaper to self-host an LLM or use an API?

At low to moderate volumes, APIs are typically cheaper because you only pay for the tokens you use. Self-hosting becomes cost-effective when sustained high throughput spreads fixed GPU costs over a large volume of tokens. The break-even point depends on your model size, utilization, and GPU rates; the calculator computes it from your own numbers.

What is the cheapest LLM API in 2026?

Among the models the calculator tracks, the lowest-priced as of its last-verified pricing include Cohere Command R7B ($0.0375 per million input tokens / $0.15 per million output), GPT-5 Nano ($0.05/$0.40), Llama 3.1 8B on Groq ($0.06/$0.06), and Mistral Small 4 ($0.10/$0.30). The best value depends on your quality requirements and use case; the calculator lets you compare within a model tier.

How do I calculate LLM API costs for my project?

Estimate your monthly token usage (1 token ≈ 0.75 words), split between input and output, then multiply by each per-million-token price. For example, 5M tokens/month at a 60/40 input/output split on GPT-5.4 ($2.50 input / $15 output per million tokens, as of the calculator's last-verified pricing): 3M input × $2.50/1M + 2M output × $15/1M = $7.50 + $30 = $37.50/month.

Related Resources

Book a free consultation