LLM API cost calculator
Monthly cost of Claude, OpenAI, Gemini and Amazon Bedrock models from your request volume and token counts, with prompt caching and batch pricing, and the same workload priced on 30 models.
System prompt, history, retrieved documents and the question.
Include reasoning or thinking tokens: they are billed as output.
The part of each prompt that repeats, such as a long system prompt, when prompt caching is on.
For work that can wait: results come back asynchronously at the vendor's batch price.
$10,800 a year · $9.00 per 1,000 requests
- Input tokens
- $400.00
- Cached input tokens
- $0.00
- Output tokens
- $500.00
- Tokens a month (in / out)
- 200M / 50M
List prices retrieved 2026-09-19. Cache-write premiums, long-context rates and tool fees are not included; see below.
Share your number
100,000 requests a month (2,000 in / 500 out tokens) on Claude Sonnet 5 ≈ $900.00/month at list price. tools.getfinops.cloud/llm-api-cost-calculator
| # | Model | Vendor | Per month |
|---|---|---|---|
| 1 | Amazon Bedrock | $14.00 | |
| 2 | Amazon Bedrock | $24.00 | |
| 3 | OpenAI | $60.00 | |
| 4 | Amazon Bedrock | $60.00 | |
| 5 | Amazon Bedrock | $96.50 | |
| 6 | OpenAI | $100.00 | |
| 7 | OpenAI | $150.00 | |
| 8 | Amazon Bedrock | $175.00 | |
| 9 | Amazon Bedrock | $180.00 | |
| 10 | Google Gemini | $185.00 |
How LLM API pricing works
Every vendor here bills per token, with separate prices for the tokens you send (input) and the tokens the model writes (output). Output costs more than input for 29 of the 30 models on this page, and reasoning or thinking tokens are billed as output.
Two discounts matter most. Prompt caching charges less for the part of a prompt the provider has already seen, such as a long system prompt or a document you ask several questions about. Batch processing trades a quick answer for a lower price: you submit requests and collect the results later.
Prices below are per million tokens, in US dollars, at the standard tier, as published on 2026-09-19.
Claude API pricing
| Model | Input | Cached input | Output | Batch input | Batch output |
|---|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $0.25 | $50 | $5 | $25 |
| Claude Opus 5 | $5 | $0.50 | $25 | $2.50 | $12.50 |
| Claude Sonnet 5 | $2 | $0.20 | $10 | $1 | $5 |
| Claude Sonnet 4.6 | $3 | $0.30 | $15 | $1.50 | $7.50 |
| Claude Haiku 4.5 | $1 | $0.10 | $5 | $0.50 | $2.50 |
OpenAI API pricing
| Model | Input | Cached input | Output | Batch input | Batch output |
|---|---|---|---|---|---|
| GPT-6 Astra | $10 | $1 | $50 | $5 | $25 |
| GPT-5.6 Sol | $4 | $0.40 | $20 | $2 | $10 |
| GPT-5.6 Terra | $2 | $0.20 | $12 | $1 | $6 |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | $0.10 | $0.60 |
| GPT-5.5 | $5 | $0.50 | $30 | $2.50 | $15 |
| GPT-5.4 mini | $0.75 | $0.075 | $4.50 | $0.375 | $2.25 |
| GPT-5 mini | $0.25 | $0.025 | $2 | $0.125 | $1 |
| GPT-4.1 | $2 | $0.50 | $8 | $1 | $4 |
| GPT-4o mini | $0.15 | $0.075 | $0.60 | $0.075 | $0.30 |
GPT-5.6 Sol: Promotional price, offered at least through November 21, 2026. GPT-5.5: Price for prompts under 272K tokens.
Gemini API pricing
| Model | Input | Cached input | Output | Batch input | Batch output |
|---|---|---|---|---|---|
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 | $0.375 | $1.875 |
| Gemini 3.5 Flash | $1.50 | $0.15 | $9 | $0.75 | $4.50 |
| Gemini 3.5 Flash-Lite | $0.30 | $0.03 | $2.50 | $0.15 | $1.25 |
| Gemini 3.1 Pro Preview | $2 | $0.20 | $12 | $1 | $6 |
| Gemini 2.5 Pro | $1.25 | $0.125 | $10 | $0.625 | $5 |
| Gemini 2.5 Flash | $0.30 | $0.03 | $2.50 | $0.15 | $1.25 |
Gemini 3.8 Flash: $1.50 input and $7.50 output from 2027-01-01. Gemini 3.1 Pro Preview: Price for prompts up to 200K tokens. Gemini 2.5 Pro: Price for prompts up to 200K tokens.
Amazon Bedrock pricing
| Model | Input | Cached input | Output | Batch input | Batch output |
|---|---|---|---|---|---|
| Amazon Nova 2 Pro | $1.375 | — | $11 | $0.6875 | $5.50 |
| Amazon Nova 2 Lite | $0.33 | $0.0825 | $2.75 | $0.165 | $1.375 |
| Amazon Nova Pro | $0.80 | $0.20 | $3.20 | $0.40 | $1.60 |
| Amazon Nova Lite | $0.06 | $0.015 | $0.24 | $0.03 | $0.12 |
| Amazon Nova Micro | $0.035 | $0.00875 | $0.14 | $0.0175 | $0.07 |
| Llama 4 Maverick 17B | $0.24 | — | $0.97 | $0.12 | $0.485 |
| Llama 3.3 70B | $0.72 | — | $0.72 | $0.36 | $0.36 |
| DeepSeek V3.2 | $0.62 | — | $1.85 | $0.31 | $0.925 |
| gpt-oss-120b | $0.15 | — | $0.60 | $0.075 | $0.30 |
| Mistral Large 3 | $0.50 | — | $1.50 | $0.25 | $0.75 |
A dash means no price is published for that model. Cross-region (global) inference and flex service are priced differently; see the AWS Bedrock pricing page.
Worked example: 100,000 requests a month
Each request sends 2,000 tokens and gets 500 back: 200 million input and 50 million output tokens a month, with no caching or batching.
- Claude Sonnet 5 (Anthropic): $900.00 a month
- GPT-5.6 Terra (OpenAI): $1,000 a month
- Gemini 3.8 Flash (Google Gemini): $337.50 a month, at its price until 2027-01-01
- Amazon Nova 2 Lite (Amazon Bedrock): $203.50 a month
The spread is wide, which is why the table above ranks every model for your own numbers. Price is only half the question: run your prompts on the cheaper candidates before you switch.
Ways to lower an LLM API bill
- Route by difficulty. Send routine requests to a small model and escalate only the ones that need a larger one.
- Cache stable prefixes. Put the system prompt, tool definitions and shared documents first and keep them identical between requests, so they can be read from the cache.
- Batch what can wait. Evaluations, backfills, classification and nightly summaries rarely need an answer in seconds.
- Cap output. Set a maximum output length and ask for short answers where they will do; output tokens cost the most.
- Send less context. Retrieve fewer, better passages instead of whole documents, and trim long conversation histories.
Frequently asked questions
How much does the Claude API cost?
Anthropic prices each model per million tokens. Claude Sonnet 5 costs $2 per million input tokens and $10 per million output tokens; Claude Haiku 4.5 costs $1 per million input tokens and $5 per million output tokens; Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens; and Claude Fable 5.1 costs $10 per million input tokens and $50 per million output tokens. Reading from the prompt cache costs a tenth of the input price (a fortieth on Claude Fable 5.1), and the Batch API halves both input and output.
How much does the OpenAI API cost?
GPT-5.6 Terra costs $2 per million input tokens and $12 per million output tokens. GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output tokens. GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens. Cached input is priced lower again, and OpenAI's batch prices are half the standard ones for the models listed here.
What does the Gemini API cost?
On the paid tier, Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens until 2027-01-01 (then $1.50 and $7.50), and Gemini 3.5 Flash-Lite costs $0.30 per million input tokens and $2.50 per million output tokens. Pro models such as Gemini 3.1 Pro Preview charge more for prompts over 200K tokens. Google also offers a free tier with lower rate limits, which this calculator does not model.
How is Amazon Bedrock priced?
Bedrock charges per input and output token for on-demand use, with separate rates for batch, for priority and flex service, and for cross-region inference. This page uses the in-region standard on-demand rate in US East (N. Virginia) for each Bedrock model listed, and for all of them the batch rate is half of it. Claude models are also sold on Bedrock; AWS publishes those prices on its Bedrock pricing page.
What is a token?
A piece of text the model reads or writes. Anthropic gives a rough guide of about 4 characters or 0.75 words of English per token. Each vendor uses its own tokenizer, so the same text can be a different number of tokens on different models: Anthropic says its tokenizer for Claude 4.7 and later produces about 30% more tokens for the same text than earlier Claude models.
How much do prompt caching and batching save?
On Claude Sonnet 5, the worked example below costs $900.00 a month. With 60% of each prompt read from the cache it falls to $684.00, and through the Batch API to $450.00. Caching helps when requests share a long, identical prefix; batching suits work that can wait for an asynchronous result.
What does this calculator leave out?
Cache-write charges (Anthropic, and OpenAI on its newest models, charge more than the input price to store a prompt in the cache; Google charges for cache storage by the hour), long-context rates, tool and web-search fees, images and audio, fine-tuning, regional or data-residency premiums, free tiers and credits, and negotiated discounts.
Prices are public list prices in US dollars, before tax, retrieved on 2026-09-19 from Anthropic API pricing, OpenAI API pricing, Gemini Developer API pricing and the AWS Price List API (AmazonBedrock). They change; check the source before committing to a number. Your bill also depends on discounts, credits and usage this calculator does not see. Bedrock prices were read from the AWS Price List on 2026-09-19.