AI API Cost Calculator

Estimate your monthly LLM spend in real time. Compare GPT-5, Claude, DeepSeek, Gemini and Llama pricing, then apply cached-input pricing and cost-based routing to see what you can save.

⚙️ Usage & Model

💰 Cost Optimization

📊 Your Estimated Cost

Monthly Cost
$0.00
$0.000000 per request
0% cheaper than running the same workload on GPT-5 at full price
Input Cost / Mo
$0.00
Output Cost / Mo
$0.00
Yearly Cost
$0.00
Savings vs GPT-5
$0.00

Prices are list prices (USD per 1M tokens) from official provider pricing. Your actual bill depends on provider discounts and usage patterns.

Model Comparison

Same workload priced across models — your selection, GPT-5 (premium) and DeepSeek Chat (budget), with your cache & routing options applied.

ModelInput $/1MOutput $/1MCost / RequestMonthly CostRelative to GPT-5

How the calculation works

LLM APIs bill per token — roughly 4 characters or 0.75 English words. Input tokens (your prompt) and output tokens (the model's reply) are priced separately, and costs scale linearly with usage:

cost = requests × (input_tokens × input_price + output_tokens × output_price) / 1,000,000

Frequently asked questions

Which models are included in the calculator?
Eight of the most-used frontier models: GPT-5, GPT-5-mini, Claude Opus 4, Claude Sonnet 4, DeepSeek R1, DeepSeek Chat, Gemini 2.5 Pro and Llama 4 — covering premium, mid-tier and budget tiers across OpenAI, Anthropic, DeepSeek, Google and Meta.
How accurate is the estimate?
It uses official per-1M-token list prices and your request count plus average token lengths, so the math is exact. The accuracy of the prediction depends on how close your real average token counts and cache hit rate are to your inputs — check your provider dashboard after a week of traffic and adjust.
What is cached-input pricing?
When a request's prefix matches a previously seen prompt, providers serve it from a cache at a large discount (often 90% off input pricing). Raising the cache hit rate in the calculator approximates workloads with stable, repetitive system prompts.
How does cost-based routing save 62%?
A model router inspects each request and sends it to the cheapest capable model — hard tasks to a frontier model, easy tasks to GPT-5-mini or DeepSeek Chat. Blended across a typical workload this cuts spend by roughly 62% (multiplier 0.38) versus single-model GPT-5 usage.
Can I use these models through DrAI?
Yes — DrAI aggregates GPT, Claude, DeepSeek, Gemini and open-source models behind one OpenAI-compatible API with per-token billing, cached-input pricing and optional smart routing. See pricing or start at DrAI Chat.

Stop guessing your AI bill

Run the same workloads on DrAI with transparent per-token pricing, cached-input discounts and smart routing — and pay only for what you use.

🌐 English