How to Estimate AI API Costs Before You Build
A simple AI API cost calculator method with real October 2026 prices for GPT-6, Claude 5.5 and Gemini, plus three worked examples showing every step of the math.
On this page
- Key takeaways
- The AI API cost formula
- Current API prices to plug into your estimate
- How to estimate AI API costs in 7 steps
- Worked example 1: a website support chatbot
- Worked example 2: generating blog drafts
- Worked example 3: summarizing long documents
- Hidden costs most API estimates miss
- How to cut the estimate before you build
- Frequently asked questions
- Next steps: build your estimate today
To estimate AI API costs, multiply the input tokens per request by the model’s input price, add the output tokens times the output price, then multiply by how many requests you expect each month. The formula is simple: cost = (input tokens / 1,000,000 x input price) + (output tokens / 1,000,000 x output price). The hard part is counting tokens honestly, including chat history, hidden reasoning tokens and tool fees.
This guide gives you a repeatable AI API cost calculator method, real October 2026 prices for OpenAI, Anthropic and Google models, and three worked examples with every step of the arithmetic shown.
Key takeaways
- Every estimate comes down to two numbers per request (input and output tokens) multiplied by the model’s price per 1M tokens and by monthly volume.
- Chatbots resend the whole conversation each turn, so a six-turn chat can use about 25 times more input tokens than the user actually typed.
- A 3,000-conversation support bot costs about $6.71 a month on GPT-6 Luna, $50.29 on Gemini 3.8 Flash and $134.10 on GPT-6.1 Sol or Claude Sonnet 5.5, before caching.
- Count tokens with free tools: OpenAI’s Tokenizer and tiktoken, and Anthropic’s free count_tokens endpoint.
- Add a buffer for reasoning tokens, retries, web search fees and long-context surcharges.

The AI API cost formula
AI APIs bill by the token, a chunk of text roughly four characters long. OpenAI’s rule of thumb is that one token is about three quarters of a word, so 100 tokens is about 75 English words. Other languages, including Hindi and other Indian languages, can split into tokens differently, so always measure real samples. If this is new, start with our explainer on what tokens are in AI.
Prices are quoted per 1 million tokens, with separate rates for input (what you send) and output (what the model writes). Output usually costs about five times more than input. The monthly formula is:
Monthly cost = requests per month x [(input tokens x input price) + (output tokens x output price)] / 1,000,000
That is the whole AI API cost calculator. Everything else in this guide is about feeding it accurate numbers.
Current API prices to plug into your estimate
These standard prices were read on the official pricing pages on 1 October 2026. All are USD per 1M tokens.
| Model | Input | Cached input | Output | Good for |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | Hardest reasoning and agent work |
| GPT-6.1 Sol | $2.00 | $0.10 | $10.00 | General production workloads |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | High-volume simple tasks |
| Claude Opus 5.5 | $4 | $0.20 | $20 | Complex writing, coding, analysis |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 | Balanced quality and cost |
| Claude Haiku 4.5 | $1 | $0.10 | $5 | Fast, cheap Claude tasks |
| Gemini 3.1 Pro Preview | $2.00 | $0.20 | $12.00 | Prompts up to 200K tokens |
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 | Price rises to $1.50 / $7.50 on 1 Jan 2027 |
Check the live OpenAI pricing page, Anthropic pricing docs and Gemini API pricing before you finalize a budget. For deeper dives, see our guides to OpenAI API pricing, Claude API pricing and Gemini API pricing.
How to estimate AI API costs in 7 steps
- Define one unit of work. Pick the thing you will be billed for repeatedly: one support conversation, one blog draft, one document summary, one classified email. Estimates fail when the unit is vague.
- Write a realistic sample prompt. Include the full system prompt, any examples, retrieved documents and the user input. Do not estimate from a one-line test prompt if production will send three pages of instructions.
- Count the input tokens. Paste the sample into OpenAI’s web Tokenizer or count it with the open-source tiktoken library. For Claude, the Messages API count_tokens endpoint is free to use and returns a figure like {“input_tokens”: 14}. Anthropic notes the count is an estimate and can differ slightly.
- Measure the output tokens. Run 10 to 20 real requests and read the usage figures in each response. Average them. Output length varies more than input, so use real data rather than a guess.
- Multiply by monthly volume. Estimate requests per day, multiply by 30, then plug everything into the formula. Model a low, expected and high case.
- Add the hidden extras. Reasoning tokens, tool calls, retries and surcharges (covered below) can add a lot. List each one explicitly.
- Apply discounts you will actually use. Caching and batch processing can cut costs sharply, but only include them if your design supports them.
Tip: Build the estimate in a spreadsheet with one row per model and columns for input tokens, output tokens, volume and price. Changing the model or the volume then updates the whole budget instantly.
Worked example 1: a website support chatbot
Chatbots are where most estimates go wrong, because every turn resends the earlier conversation. Here are the assumptions:
- System prompt (instructions and FAQs): 1,500 tokens.
- Each user message: 100 tokens. Each bot reply: 250 tokens.
- Average conversation: 6 turns.
- Volume: 3,000 conversations per month (about 100 a day).
Step 1: input tokens per conversation
Turn 1 sends the system prompt plus the first message: 1,500 + 100 = 1,600 tokens. Each later turn also resends the previous exchange of 350 tokens (100 + 250). So turn 2 sends 1,950 tokens, turn 3 sends 2,300, and turn 6 sends 3,350.
Adding all six turns: (6 x 1,600) + 350 x (0 + 1 + 2 + 3 + 4 + 5) = 9,600 + 5,250 = 14,850 input tokens per conversation. The user only typed 600 of those.
Step 2: output tokens per conversation
6 replies x 250 tokens = 1,500 output tokens.
Step 3: monthly totals
Input: 14,850 x 3,000 = 44,550,000 tokens (44.55M). Output: 1,500 x 3,000 = 4,500,000 tokens (4.5M).
Step 4: price it
| Model | Input cost | Output cost | Monthly total |
|---|---|---|---|
| GPT-6 Luna | 44.55 x $0.10 = $4.46 | 4.5 x $0.50 = $2.25 | $6.71 |
| Gemini 3.8 Flash | 44.55 x $0.75 = $33.41 | 4.5 x $3.75 = $16.88 | $50.29 |
| Claude Haiku 4.5 | 44.55 x $1 = $44.55 | 4.5 x $5 = $22.50 | $67.05 |
| GPT-6.1 Sol | 44.55 x $2 = $89.10 | 4.5 x $10 = $45.00 | $134.10 |
| Claude Sonnet 5.5 | 44.55 x $2 = $89.10 | 4.5 x $10 = $45.00 | $134.10 |
Step 5: apply prompt caching
The 1,500-token system prompt is sent on every turn: 6 x 1,500 x 3,000 = 27M tokens a month. On Claude Sonnet 5.5, those cost 27 x $2 = $54 at the normal rate, or 27 x $0.20 = $5.40 as cache hits. If traffic is steady enough to keep the cache warm, the monthly bill falls from $134.10 to about $134.10 minus $54 plus $5.40 = $85.50, before small cache-write charges. Our guide to prompt caching explains cache lifetimes and write fees.
Building this bot? Our walkthrough on how to add an AI chatbot to your website covers the setup, and our roundup of AI customer support tools compares ready-made options if you would rather pay per seat.
Worked example 2: generating blog drafts
A content team wants 100 long-form drafts a month. Each request sends a 1,000-token brief and gets back a 2,000-word draft. Using the 0.75 words per token rule, 2,000 words is about 2,000 / 0.75 = 2,667 output tokens.
- Monthly input: 100 x 1,000 = 100,000 tokens (0.1M).
- Monthly output: 100 x 2,667 = 266,700 tokens (about 0.267M).
On Claude Opus 5.5: 0.1 x $4 = $0.40 input, plus 0.267 x $20 = $5.33 output, so about $5.73 a month. On GPT-6 Astra: 0.1 x $10 = $1.00, plus 0.267 x $50 = $13.33, so about $14.33. On GPT-6.1 Sol: 0.1 x $2 = $0.20, plus 0.267 x $10 = $2.67, so about $2.87.
Content generation is cheap per piece, even on premium models. The real cost is editing time. If you plan to write with AI, our guide on how to write a blog post with AI covers the workflow that keeps drafts usable.
Worked example 3: summarizing long documents
A legal or research team summarizes 500 contracts a month. Each contract is about 60,000 tokens and each summary is 800 tokens.
- Input: 500 x 60,000 = 30M tokens. Output: 500 x 800 = 0.4M tokens.
- Claude Sonnet 5.5: 30 x $2 = $60.00, plus 0.4 x $10 = $4.00, so $64.00.
- Gemini 3.8 Flash: 30 x $0.75 = $22.50, plus 0.4 x $3.75 = $1.50, so $24.00.
- Same job in batch mode (50% off): $32.00 on Sonnet 5.5, $12.00 on Gemini 3.8 Flash.
Input dominates this budget, which is typical for summarization. Since nobody needs the summaries within seconds, this is a perfect fit for the techniques in our batch API guide.
Hidden costs most API estimates miss
Reasoning tokens
Reasoning models think before they answer, and OpenAI confirms those internal tokens count toward output usage and billing. You cannot see them in the reply, but they appear in the usage data. Measure them on real samples. For planning, it is sensible to model a high case where output is two or three times the visible answer, then refine once they have data.
Long-context surcharges
On OpenAI’s GPT-6 and GPT-5.6 models, any request with more than 272K input tokens is billed at 2x the input rate and 1.5x the output rate for the full request. Gemini 3.1 Pro Preview jumps from $2 / $12 to $4 / $18 above 200K tokens. Claude 4.6 and later models have no long-context surcharge. Our guide to the context window explains how to keep prompts lean.
Tools and grounding
- OpenAI web search: $10 per 1,000 calls, plus search content tokens at the model’s rate.
- Anthropic web search: $10 per 1,000 searches. Code execution: $0.05 per hour after 50 free hours per day.
- Gemini Grounding with Google Search: 5,000 free requests per month across Gemini 3.x models, then $14 per 1,000.
Regional processing
OpenAI’s data residency adds 10% for models released on or after 5 March 2026. Anthropic’s US-only inference costs 1.1x on all token categories for 4.6 and later models. If your client requires either, multiply the estimate by 1.1.
Retries and failed calls
Timeouts, malformed outputs and retries all burn tokens. Add a buffer of 10% to 20% until real logs show your actual rate. That buffer is an editorial rule of thumb, not a vendor figure, so replace it with measured data as soon as you can.
Watch out: Chat subscriptions do not cover API usage. A ChatGPT Plus or Claude Pro plan gives you nothing on the developer platform, so budget the API separately.
How to cut the estimate before you build
- Pick the smallest model that passes your tests. In example 1, GPT-6 Luna costs about one twentieth of GPT-6.1 Sol. Our list of the cheapest AI model for each task helps you shortlist.
- Trim conversation history. Summarize older turns or keep only the last few. In example 1, history added 5,250 tokens per conversation.
- Cap output length. Output costs about five times input, so ask for concise answers and set a maximum.
- Cache shared prefixes. Put static instructions first so they hit the cache.
- Batch background work. Anything that can wait gets 50% off.
Our guide on how to reduce token usage in ChatGPT and Claude covers more techniques, and if you want to start for free, Gemini Flash and Flash-Lite models are free of charge on Google’s API free tier, with limits shown in AI Studio.
Set spending guardrails
Google’s Gemini API caps Tier 1 billing accounts at $250 until you have paid $100 and waited three days to reach Tier 2. Whatever provider you use, set budget alerts in the console so a bug cannot run up a surprise bill.
Indian teams are billed in USD for these APIs. Convert your final estimate at the rate on your card statement, and allow for any foreign currency markup your bank adds.
Frequently asked questions
Is there a free AI API cost calculator?
You do not need a dedicated tool. Count tokens with OpenAI’s free Tokenizer or tiktoken library, or Anthropic’s free count_tokens endpoint, then use the formula: requests x (input tokens x input price + output tokens x output price) / 1,000,000. A simple spreadsheet with one row per model works as a reliable calculator and is easy to update when prices change.
How many tokens is 1,000 words?
Using OpenAI’s rule of thumb that one token is about three quarters of a word, 1,000 English words is roughly 1,333 tokens. Code, tables, unusual formatting and non-English languages can use more tokens per word, so count a real sample before you rely on this figure in a budget.
How much does it cost to run an AI chatbot per month?
It depends on model, traffic and conversation length. In our example of 3,000 six-turn conversations a month, costs ranged from about $6.71 on GPT-6 Luna to $134.10 on GPT-6.1 Sol or Claude Sonnet 5.5, before caching. Caching the system prompt cut the Sonnet 5.5 figure to about $85.50.
Why is my real API bill higher than my estimate?
The usual causes are resent chat history, reasoning tokens billed as output, longer replies than expected, retries, tool fees such as web search, and long-context surcharges on very large prompts. Compare the usage fields in your API responses with your assumptions, then update the spreadsheet with measured averages.
Is the OpenAI API cheaper than ChatGPT Plus?
For light personal use, often yes. ChatGPT Plus is a flat $20 a month, while 100 long drafts on GPT-6.1 Sol cost about $2.87 in API fees in our example. But the API has no chat interface, memory or built-in tools, so you need an app or code to use it. Heavy users may still prefer the predictable subscription.
Next steps: build your estimate today
Pick one unit of work, count its tokens on three real samples, and price it on two or three models in a spreadsheet. Then add reasoning, tools and a buffer, and decide which discounts your design supports. Once the product is live, compare the estimate with real usage every month and adjust. Our guide on how to measure AI ROI shows how to turn those numbers into a business case. API prices change frequently, so confirm current rates on each vendor’s pricing page before you commit to a budget.
Pricing and features are checked at the time of writing and can change. Some links may be affiliate links, which never affect our verdicts.