OpenAI API Pricing: Cost Per Million Tokens for Every Model
Every OpenAI API rate for September 2026, from GPT-6 Astra to GPT-6 Luna, with worked cost examples, hidden surcharges and proven ways to cut your bill.
On this page
- Key takeaways
- OpenAI API pricing table: cost per million tokens (September 2026)
- Each OpenAI model explained: who should use it
- What does the OpenAI API really cost? Worked examples
- Hidden costs in OpenAI API pricing
- Image, realtime and audio API pricing
- How to save money on the OpenAI API
- OpenAI API vs Claude and Gemini on price
- Which OpenAI model should you pick?
- Frequently asked questions
- Final verdict on OpenAI API pricing
OpenAI API pricing is charged per million tokens, with separate rates for input, cached input and output. As of September 2026, the flagship GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens, the balanced GPT-6.1 Sol costs $2 and $10, and the budget GPT-6 Luna costs just $0.10 and $0.50.
That spread is huge: Astra output costs 100 times more than Luna output. So the real question is not “how much does the OpenAI API cost” but “which model, which tier and which discounts fit my workload”. This guide lists every current rate, shows the arithmetic on realistic examples, and explains the surcharges and savings most developers miss.
Key takeaways
- GPT-6 Astra is $10 input / $50 output per 1M tokens, GPT-6.1 Sol is $2 / $10, and GPT-6 Luna is $0.10 / $0.50.
- The GPT-5.6 family is still sold, and GPT-5.6 Sol is on promotional pricing ($4 / $20) at least through November 21, 2026.
- Batch and Flex cut token prices by 50%, and cached input is 90% to 95% cheaper than fresh input.
- Requests over 272K input tokens are billed at 2x input and 1.5x output for the whole request, so long prompts get expensive fast.
- API usage is billed separately from ChatGPT Plus or Pro. A ChatGPT subscription does not include API credits.
OpenAI API pricing table: cost per million tokens (September 2026)
All prices below are the Standard tier, in USD per 1 million tokens, taken from OpenAI’s model pages. If tokens are a new idea for you, our plain English guide to AI tokens explains them in five minutes. A rough rule: 1 million tokens is around 750,000 English words.
| Model | Input | Cached input | Output | Context window | Best for |
|---|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | 1,050,000 | Hardest reasoning, agents, high-stakes output |
| GPT-6.1 Sol | $2.00 | $0.10 | $10.00 | 1,050,000 | Default choice for most production apps |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | 1,050,000 | Classification, extraction, high-volume chat |
| GPT-5.6 Sol (promo) | $4.00 | $0.40 | $20.00 | 1,050,000 | Existing GPT-5.6 integrations |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | 1,050,000 | Mid-tier legacy workloads |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | 1,050,000 | Cheap legacy workloads |
Every model in the table supports up to 128,000 output tokens per response. GPT-6 Astra has a knowledge cutoff of April 30, 2026 and GPT-6 Luna of May 18, 2026, while the GPT-5.6 models stop at February 16, 2026. You can confirm the numbers on the official OpenAI API pricing page.
Note: OpenAI no longer uses “mini” and “nano” names for its new families. The tiers are now named Astra, Sol, Terra and Luna, and Luna is always the budget option.

Each OpenAI model explained: who should use it
GPT-6 Astra: the flagship ($10 / $50)
Astra is OpenAI’s most capable model and supports five reasoning effort levels: low, medium, high, xhigh and max. Higher effort means the model “thinks” longer, and those hidden reasoning tokens are billed as output. That matters because Astra’s $50 output rate is where the bill grows.
Use Astra when a wrong answer is costly: complex code changes, legal or financial analysis, multi-step agents that must plan well. Skip it for chat widgets, summaries and routine content, where Sol gives you most of the quality at a fifth of the price.
GPT-6.1 Sol: the sensible default ($2 / $10)
Sol is the model most teams should start with. It is five times cheaper than Astra on both input and output, and its cached input rate of $0.10 is a 95% discount, the steepest caching discount in the OpenAI lineup. For apps with a long, fixed system prompt, that makes Sol very cheap to run.
It fits customer support bots, content drafting, RAG (retrieval augmented generation, where you feed the model your own documents), and most coding assistants.
GPT-6 Luna: the budget workhorse ($0.10 / $0.50)
Luna is priced for volume. At $0.10 per million input tokens you could send 10 million tokens of text for $1. It is ideal for tagging, routing, sentiment analysis, simple extraction and first-pass filtering before a stronger model sees the hard cases. Our guide on picking the cheapest AI model for each task shows how to build this kind of routing.
GPT-5.6 Sol, Terra and Luna: still on sale
The older GPT-5.6 family is still available. GPT-5.6 Sol is currently on promotional pricing of $4 / $0.40 / $20, which OpenAI says runs “at least through November 21, 2026”. Many third-party sites still quote its older list price of $5 input and $30 output, so do not be surprised if you see different numbers elsewhere.
For new projects there is little reason to pick GPT-5.6. GPT-6.1 Sol costs half as much as GPT-5.6 Sol and GPT-6 Luna costs half as much as GPT-5.6 Luna. Keep GPT-5.6 only if you have tested prompts that depend on its behavior.
Watch out: Older models such as GPT-5.5, GPT-5.4 and o3 are still quoted on third-party sites, but we could not confirm their current rates on an official OpenAI page. Check the live pricing page before you build on them.
What does the OpenAI API really cost? Worked examples
Per-token prices are hard to picture, so here are three realistic workloads with the arithmetic shown. For a deeper method, see how to estimate AI API costs before you build.
Example 1: a customer support chatbot
Say your bot handles 1,000 conversations a day. Each conversation sends about 2,000 input tokens (system prompt, history, the question) and gets back 500 output tokens. Per day that is 2 million input tokens and 0.5 million output tokens.
- GPT-6 Astra: (2 x $10) + (0.5 x $50) = $20 + $25 = $45 per day, about $1,350 a month.
- GPT-6.1 Sol: (2 x $2) + (0.5 x $10) = $4 + $5 = $9 per day, about $270 a month.
- GPT-6 Luna: (2 x $0.10) + (0.5 x $0.50) = $0.20 + $0.25 = $0.45 per day, about $13.50 a month.
If you are planning this exact build, our walkthrough on adding an AI chatbot to your website covers the setup side.
Example 2: the same bot with prompt caching
Now assume 1,500 of those 2,000 input tokens are an identical system prompt and product FAQ at the start of every request. That prefix is above the 1,024-token caching minimum, so OpenAI caches it automatically. On GPT-6.1 Sol, cached reads cost $0.10 per million instead of $2.
- Cached input: 1.5M tokens x $0.10 = $0.15
- Fresh input: 0.5M tokens x $2 = $1.00
- Output: 0.5M tokens x $10 = $5.00
- Total: about $6.15 per day instead of $9, a saving of roughly 32%.
The first request that writes the cache costs a little more (see the next section), but with steady traffic the cache stays warm and the savings dominate. Output is now the biggest cost, which is typical once caching is on.
Example 3: a nightly batch of document summaries
You summarize 5,000 reports overnight, each 4,000 tokens in and 400 tokens out. That is 20M input and 2M output tokens. On GPT-6.1 Sol Standard: (20 x $2) + (2 x $10) = $40 + $20 = $60. Through the Batch API at 50% off, the same job costs $30, and results arrive within 24 hours.
Hidden costs in OpenAI API pricing
The long-context surcharge above 272K tokens
All GPT-6 and GPT-5.6 models accept over a million tokens of context, but any request with more than 272K input tokens is billed at 2x the input and cache rates and 1.5x the output rate, for the full request, not just the part above the threshold.
Example on GPT-6.1 Sol: a 300,000-token prompt with a 5,000-token answer. Normally that would be (0.3 x $2) + (0.005 x $10) = $0.60 + $0.05 = $0.65. With the surcharge it becomes (0.3 x $4) + (0.005 x $15) = $1.20 + $0.075 = about $1.28, nearly double. If you can trim the prompt to under 272K, or split it, do it. Our explainer on how context windows work covers chunking strategies.
Cache write fees
On GPT-5.6 and later models, writing a prompt into the cache costs 1.25x the normal input rate. Cached entries stay valid for 30 minutes after the last write or reuse. If your traffic is sparse and the cache expires between requests, you pay the write premium without getting the reads, so caching helps busy apps far more than occasional scripts.
Reasoning tokens
When Astra runs at high, xhigh or max effort, it generates internal reasoning that is billed at the output rate. A short visible answer can hide many more billed tokens. Start at low or medium effort and raise it only for tasks that fail.
Tools, data residency and Fast mode
- Web search tool: $10 per 1,000 calls, plus the search content tokens billed at the model’s normal rate.
- Data residency: regional processing adds 10% for models released on or after March 5, 2026.
- Fast (priority) mode: costs 2x Standard. OpenAI describes it as “twice as much for double the speed”.
Image, realtime and audio API pricing
Beyond text, OpenAI prices its media models by token type or by minute. These rates were read on OpenAI’s pricing pages; the image model names in particular are worth double-checking on the live page.
| Model | Pricing (Standard tier) |
|---|---|
| gpt-image-2 | Image tokens: $8 input, $2 cached, $30 output. Text tokens: $5 input, $1.25 cached. |
| gpt-realtime-2.1 | Text: $4 / $0.40 / $24. Audio: $32 / $0.40 / $64. Image input $5. |
| gpt-realtime-2.1-mini | Text: $0.60 / $0.06 / $2.40. Audio: $10 / $0.30 / $20. |
| GPT-Transcribe | $0.0045 per minute |
| GPT-Live-Transcribe | $0.017 per minute |
| GPT-Realtime-Translate | $0.034 per minute |
OpenAI does not publish a simple per-image price for gpt-image-2. Cost depends on quality and size, and OpenAI points to a calculator in its image generation guide. For comparing image tools on price and quality, see Midjourney vs ChatGPT image generation.
Note: The Sora video API was discontinued on September 24, 2026, after the Sora app closed in April. There is no OpenAI video generation API to budget for right now.
How to save money on the OpenAI API
- Route by difficulty. Send easy requests to GPT-6 Luna and only escalate hard ones to Sol or Astra. Even if 20% of traffic needs Sol, your blended cost drops sharply.
- Put static content first. Caching matches on the prefix, so keep your system prompt, instructions and reference docs at the very start and put user-specific content at the end. Our prompt caching guide shows the right prompt layout, and OpenAI’s prompt caching docs list the exact rules.
- Use Batch or Flex for anything not urgent. Both are half the Standard price. Batch jobs finish within 24 hours and accept up to 50,000 requests or 200 MB per batch. Read our Batch API guide for setup tips, or the official Batch documentation.
- Cap output length. Output costs five to six times more than input on every current OpenAI text model. Ask for bullet points, set a max output token limit, and avoid “explain your reasoning” unless you need it.
- Stay under 272K input tokens. Retrieval that sends only relevant chunks beats stuffing whole documents into the prompt.
- Trim conversation history. Summarize old turns instead of resending them. Our tips on reducing token usage apply directly to API apps.
Save money: Combining techniques stacks the savings. A Batch job on GPT-6.1 Sol with a cached prefix pays half price on output and a tiny fraction of the normal rate on the cached input. Check the pricing page for exactly how Batch and cache discounts combine for your model.
OpenAI API vs Claude and Gemini on price
OpenAI is not the only option, and rates differ by model tier. Here is a quick comparison of mid-tier and flagship models from the three big providers, all per 1M tokens on standard pricing.
| Provider and model | Input | Output |
|---|---|---|
| OpenAI GPT-6.1 Sol | $2.00 | $10.00 |
| Anthropic Claude Sonnet 5.5 | $2.00 | $10.00 |
| Google Gemini 3.1 Pro (preview, up to 200K prompt) | $2.00 | $12.00 |
| OpenAI GPT-6 Astra | $10.00 | $50.00 |
| Anthropic Claude Opus 5.5 | $4.00 | $20.00 |
| OpenAI GPT-6 Luna | $0.10 | $0.50 |
At the mid tier, the three are close on list price, so quality on your own prompts should decide. At the top end, Claude Opus 5.5 is well below GPT-6 Astra on price, and at the budget end GPT-6 Luna is one of the cheapest models from a major lab. For the full detail, read our Claude API pricing breakdown and our Gemini API pricing guide. If cost is your top priority, DeepSeek vs ChatGPT covers an even cheaper alternative.
Which OpenAI model should you pick?
- Solo developer or student prototyping: start with GPT-6 Luna. Your experiments will cost cents, and you can upgrade later.
- Startup shipping a chatbot or writing tool: GPT-6.1 Sol with a cached system prompt. It is the best mix of quality and cost.
- Agency running bulk content or data jobs: GPT-6.1 Sol or Luna through the Batch API for the 50% discount.
- Team building coding agents or complex analysis: GPT-6 Astra at low or medium effort, with Sol handling simpler sub-tasks.
- Non-developers who just want to chat: skip the API. A subscription is simpler, and our ChatGPT pricing guide compares every plan.
Frequently asked questions
Is there a free tier for the OpenAI API?
We could not confirm a free API tier on any official OpenAI page as of September 2026. Some third-party sources say new accounts may receive a small amount of trial credit, but you should not count on it. Plan to add a payment method, set a monthly spending limit in your account, and start with GPT-6 Luna so early testing costs only cents.
Does ChatGPT Plus include OpenAI API access?
No. OpenAI bills API usage separately from ChatGPT subscriptions, so paying $20 a month for Plus or more for Pro gives you no API credits. The API is pay as you go, charged per million tokens. If you only chat in the browser, a subscription is enough; if you build apps or automations, you need a separate API account with billing.
What is the cheapest OpenAI API model?
GPT-6 Luna is the cheapest current text model at $0.10 per million input tokens, $0.01 cached and $0.50 output. It still offers the same 1,050,000-token context window as the flagship. Through the Batch API it drops to half those rates. It suits classification, extraction, routing and simple chat, but use Sol or Astra for complex reasoning.
How much does GPT-6 Astra cost per request?
It depends on tokens. A request with 2,000 input tokens and 500 output tokens costs (0.002 x $10) + (0.0005 x $50), which is $0.02 plus $0.025, or about $0.045. Higher reasoning effort adds hidden output tokens, and prompts over 272K input tokens trigger a surcharge, so real costs can be higher.
How much does the OpenAI API cost in India?
OpenAI prices the API in US dollars worldwide, and we found no separate rupee price list for API tokens. Indian developers pay the same USD rates, charged to an international card, and the final rupee amount depends on the exchange rate, bank fees and any applicable taxes. Setting a USD budget cap helps avoid surprises.
Final verdict on OpenAI API pricing
OpenAI’s 2026 lineup gives you a clear ladder: Luna for volume, Sol for most real products, and Astra for the hardest problems. The biggest cost levers are not the list prices but how you use them: route easy work to cheaper models, cache your prompt prefix, batch anything that can wait, and keep prompts under the 272K surcharge line. Do that and even a busy app can run on a modest monthly bill.
API rates and promotions change often, so always confirm current prices on OpenAI’s official pricing page before you commit to a budget.
Pricing and features are checked at the time of writing and can change. Some links may be affiliate links, which never affect our verdicts.