Context Window Explained: How Much Text Can AI Handle?
What a context window is, official 2026 sizes for GPT-6, Claude Opus 5.5 and Gemini 3.8 Flash, what long prompts really cost, and how to work within any limit.
On this page
- Key takeaways
- What is a context window, in plain English?
- Context window sizes in 2026: API models
- Context windows in ChatGPT, Claude and Gemini apps
- What long context really costs (worked examples)
- Is a bigger context window always better?
- 7 ways to work within any context window
- Frequently asked questions
- The bottom line
A context window is the maximum amount of text an AI model can read and keep in mind at once, measured in tokens. It covers everything in the conversation: your instructions, uploaded files, earlier messages and the model’s reply. In 2026 the leading API models handle around 1 million tokens, roughly 750,000 English words, but the apps you chat in often give you less, and bigger inputs cost more.
This guide explains how context windows work, lists the official sizes for ChatGPT, Claude, Gemini and other models as of 1 October 2026, shows what long context really costs with worked examples, and gives practical ways to work within the limit.
Key takeaways
- A context window is measured in tokens; in English, 1 token is about 4 characters or three quarters of a word.
- GPT-6 and GPT-5.6 API models offer 1,050,000 tokens, Claude Opus 5.5 and Sonnet 5.5 offer 1M, and Gemini 3.8 Flash accepts 1,048,576 input tokens.
- In apps, Claude’s paid plans default to 200K tokens and reach 1M on the newest models; OpenAI does not publish ChatGPT’s per-plan context size.
- Long context costs more: OpenAI doubles input prices above 272K tokens, Gemini 3.1 Pro charges more above 200K, and Claude 4.6 and later models have no surcharge.
- Bigger is not always better. Sending only what the model needs is cheaper and often more accurate.

What is a context window, in plain English?
Think of the context window as the model’s working memory for one conversation. Everything the model can “see” when it writes the next reply must fit inside it. That includes the system instructions, your messages, any documents you upload, the model’s own earlier answers, and the new reply it is about to write.
The window is measured in tokens, the small chunks of text models read. OpenAI’s rule of thumb for English is that “1 token is approximately 4 characters” and “100 tokens are approximately 75 words”. Other languages, including Hindi and other Indian languages, can split into tokens differently, often using more tokens for the same meaning. Our guide to what tokens are in AI explains this in detail.
Context window vs max output
Models have two separate limits. The context window is the total the model can handle. The max output limit is how much it can write in a single reply. GPT-6 Astra, for example, has a 1,050,000-token context window but a 128,000-token max output. Reasoning models also produce hidden “reasoning tokens” before the visible answer, and OpenAI says these count toward output usage and billing.
What happens when you go over
In the API, a request that exceeds the window is rejected or must be trimmed. In chat apps, older parts of the conversation are dropped or summarized. Claude, for example, summarizes earlier messages to make room when a chat nears the limit and code execution is turned on. That keeps the chat going, but fine details from the start of the conversation may be lost.
Context window sizes in 2026: API models
These are the official figures from each vendor’s model documentation, checked in late September and early October 2026.
| Model | Context window | Max output | Rough English words |
|---|---|---|---|
| GPT-6 Astra, GPT-6.1 Sol, GPT-6 Luna | 1,050,000 tokens | 128,000 | About 787,000 |
| GPT-5.6 Sol, Terra, Luna | 1,050,000 tokens | 128,000 | About 787,000 |
| Claude Opus 5.5, Sonnet 5.5, Fable 5.1 | 1M tokens | 128,000 | About 750,000 |
| Claude Haiku 4.5 | 200K tokens | 64,000 | About 150,000 |
| Gemini 3.8 Flash | 1,048,576 input tokens | 65,536 | About 786,000 |
| DeepSeek Flash and V4 Pro | 1M tokens | 384,000 | About 750,000 |
Word estimates use OpenAI’s three-quarters-of-a-word rule, so 1,050,000 x 0.75 = 787,500. If a printed page holds about 500 words, 750,000 words is roughly 1,500 pages. Sources include the OpenAI API pricing and models docs and Anthropic’s models overview.
Open models you run yourself are usually smaller. Gemma 4 on Ollama, for instance, lists 128K tokens for its edge variants and 256K for the 12B, 26B and 31B versions. See our guide on how to run an LLM locally if privacy matters more than size.
Context windows in ChatGPT, Claude and Gemini apps
The context you get in a chat app is often smaller than the API maximum, and it can depend on your plan.
| App and plan | Context window | Status |
|---|---|---|
| Claude paid plans (default) | 200K tokens | Official |
| Claude: Fable 5.1, Opus 5.5, Opus 5, Sonnet 5.5, Sonnet 5 | 1M tokens | Official |
| Claude: Fable 5, Opus 4.8, 4.7, 4.6, Sonnet 4.6 | 500K tokens | Official |
| Claude Free | “Up to 1M tokens (varies by model)” | Official wording, no per-model detail |
| Google AI Pro (Gemini app) | 1M tokens | Official |
| ChatGPT (all plans) | Not published by OpenAI | Third-party estimates only |
OpenAI does not publish a per-plan context window for ChatGPT, and none of the official help pages we checked on 1 October 2026 list one. A third-party table estimates around 27K tokens on Free, 54K for instant models and 256K for reasoning models on Plus, and more on Pro, but treat those as unconfirmed. The Pro plan page only promises “maximum memory and context”. If context size is critical, test with your own documents or use the API. Our Claude review covers long-document work in more depth.
Watch out: In apps, a bigger context also uses more of your usage allowance, because the model processes the whole conversation on every reply. See our guides to Claude usage limits and ChatGPT limits by plan.
What long context really costs (worked examples)
In the API you pay for every input token, every time. That makes context the biggest cost driver for document-heavy apps. Pricing rules also change above certain sizes.
Example 1: one 300,000-token document, 2,000-token answer
| Model | Rule | Arithmetic | Cost |
|---|---|---|---|
| Claude Sonnet 5.5 | No long-context surcharge | 0.3 x $2 + 0.002 x $10 | $0.62 |
| GPT-6.1 Sol | Over 272K: 2x input, 1.5x output | 0.3 x $4 + 0.002 x $15 | $1.23 |
| Gemini 3.1 Pro (preview) | Over 200K: $4 input, $18 output | 0.3 x $4 + 0.002 x $18 | About $1.24 |
OpenAI’s surcharge applies to the full request once input passes 272K tokens on GPT-6 and GPT-5.6 models. So the same document on GPT-6.1 Sol costs about twice as much in one piece as it would below the threshold. Splitting it into two 150,000-token requests, each with a 2,000-token answer, costs 2 x (0.15 x $2 + 0.002 x $10) = 2 x $0.32 = $0.64.
Example 2: asking ten questions about the same document
Resending a 300,000-token document ten times on Claude Sonnet 5.5 costs 10 x 0.3 x $2 = $6.00 in input alone. With prompt caching, you pay once to write the cache (5-minute write at $2.50 per 1M: 0.3 x $2.50 = $0.75), then cache hits at $0.20 per 1M for the other nine (9 x 0.3 x $0.20 = $0.54). Total input: $1.29, about 78 percent less. Our prompt caching guide shows how to set this up, and how to estimate AI API costs helps you model your own workload.
Save money: For bulk document jobs that can wait, the Batch APIs from OpenAI, Anthropic and Google cut token prices by 50 percent, and Anthropic says caching and batch discounts can stack. See our Batch API guide.
Is a bigger context window always better?
No. A large window is a capability, not a target. Three reasons to send less:
- Cost. As the examples show, input tokens add up fast, and some vendors charge more above a threshold.
- Speed. More input means longer processing before the first word of the answer appears.
- Accuracy. In everyday use, models are more reliable when the relevant passage is easy to find. Details buried in a huge, mostly irrelevant input are easier to miss than the same details in a focused excerpt.
Long context shines when the task truly needs the whole thing at once: comparing clauses across a full contract, tracing a bug across many files, or synthesizing a long research report. For question-and-answer over a big library of documents, retrieval (searching for the relevant chunks first and sending only those) is usually cheaper. Tools like NotebookLM do this for you; see how to use NotebookLM for study and research.
7 ways to work within any context window
- Send only the relevant part. Cut appendices, boilerplate, duplicate pages and images the model does not need.
- Start fresh chats for new tasks. Old messages fill the window and are reprocessed every turn.
- Summarize before you continue. Ask the model for a tight summary of decisions so far, then start a new chat with just that summary.
- Put stable material first. Instructions and reference documents at the start of the prompt make caching possible in the API, since caches match on an identical prefix.
- Chunk very large jobs. Process a long book chapter by chapter, then combine the summaries in a final pass.
- Use projects for reusable files. Claude caches project content so it counts less against your limits when reused.
- Count tokens before you send. OpenAI offers a web tokenizer and the tiktoken library, and Anthropic’s token counting endpoint is free to use. More tactics are in our guide to reducing token usage.
Frequently asked questions
What is a context window in AI?
A context window is the maximum number of tokens an AI model can process at once, including your instructions, uploaded files, the conversation history and the model’s reply. In English, a token is roughly three quarters of a word. Anything outside the window is invisible to the model, so very long chats get trimmed or summarized as they grow.
Which AI has the largest context window?
Among the major vendors in October 2026, OpenAI’s GPT-6 and GPT-5.6 API models list 1,050,000 tokens, Gemini 3.8 Flash accepts 1,048,576 input tokens, and Claude Opus 5.5, Sonnet 5.5 and Fable 5.1 offer 1M. They are close enough that pricing, output limits and accuracy matter more than the headline number.
What is ChatGPT’s context window?
OpenAI does not publish a context window for ChatGPT plans. The API models behind it support up to 1,050,000 tokens, but the app may give you less depending on plan and model. Third-party estimates exist, such as about 54K tokens for instant models on Plus, but they are unconfirmed. For guaranteed size, use the API.
How many words is 1 million tokens?
Using OpenAI’s rule of thumb that 100 tokens is about 75 English words, 1 million tokens is roughly 750,000 words. At around 500 words per page, that is about 1,500 pages. Other languages, code and numbers often use more tokens per word, so the real figure varies with your content.
Does a bigger context window cost more?
Yes, in two ways. In the API you pay for every input token on every request, so a large document resent many times gets expensive. Some vendors also charge more above a threshold: OpenAI doubles input rates above 272K tokens on GPT-6 and GPT-5.6, and Gemini 3.1 Pro raises prices above 200K. Claude 4.6 and later models have no surcharge.
The bottom line
Context windows have grown to around 1 million tokens across the top API models, which is enough for most real documents and many codebases. The limit you actually feel is usually your app plan, your usage allowance or your budget, not the model. Send focused inputs, cache what you reuse, and choose Claude 4.6 or later when you truly need very long prompts without a surcharge. Vendors update models, context sizes and prices often, so confirm current figures on the official pricing and documentation pages before you build.
Related reading: Claude API pricing and OpenAI API pricing.
Pricing and features are checked at the time of writing and can change. Some links may be affiliate links, which never affect our verdicts.