Skip to content
DigitalTatva
Token Optimization

What Are Tokens in AI? A Plain English Guide

Tokens are the chunks of text AI models read and write. Learn what they are, how they drive usage limits and API bills, and how to count and price them with real examples.

What Are Tokens in AI? A Plain English Guide
On this page
  1. Key takeaways
  2. What is a token in AI? The plain English version
  3. Why tokens matter even if you never touch an API
  4. Input tokens, output tokens and the other types
  5. How much do tokens cost? Real prices per million
  6. Worked examples: turning tokens into dollars
  7. How to count tokens yourself
  8. Common myths about AI tokens
  9. Practical ways to use fewer tokens
  10. Frequently asked questions
  11. Next steps: put tokens to work

Tokens are the small chunks of text that AI models read and write. A token is usually a piece of a word: OpenAI’s rule of thumb is that 1 token is about 4 characters of English, or about three-quarters of a word, so 100 tokens is roughly 75 words. Every AI tool counts tokens to decide how much text it can handle at once, how fast you hit your usage limit, and how much an API call costs.

If you use ChatGPT, Claude or Gemini in the browser, tokens explain why long chats slow down and why you run out of messages. If you build with an API, tokens are literally your bill. This guide explains tokens in plain English first, then shows the real prices and the arithmetic behind them.

Key takeaways

  • A token is a chunk of text, roughly 4 characters or three-quarters of an English word. 1,000 words is about 1,333 tokens.
  • AI APIs charge per million tokens, with separate prices for input (what you send) and output (what the model writes). Output is usually 5 times more expensive.
  • Prices range widely: GPT-6 Luna output costs $0.50 per million tokens, while GPT-6 Astra output costs $50.
  • Chat apps do not bill per token, but tokens still drive usage limits, which is why long conversations and big files drain your allowance faster.
  • Reasoning models produce hidden “thinking” tokens that are billed as output, even though you never see them.

What is a token in AI? The plain English version

Think of a token as a Lego brick of language. Before a model reads your message, a tool called a tokenizer breaks the text into these bricks. Common short words like “the” or “cat” are often a single token. Longer or rarer words get split into several pieces, so a word like “unbelievable” might become two or three tokens.

Spaces, punctuation and numbers count too. A comma can be its own token, and a long number may be split into several. That is why the word count of your text and its token count never match exactly.

The model also writes its answer one token at a time. When you watch ChatGPT or Claude “type” a reply, you are watching it produce tokens in sequence. Each new token is a prediction of what should come next, based on all the tokens before it.

Quick conversions you can use

These are OpenAI’s published rules of thumb for English. They are estimates, because sentence and word lengths vary.

Text Approximate tokens
1 token About 4 characters, or three-quarters of a word
75 words (a short paragraph) About 100 tokens
750 words (a short blog post) About 1,000 tokens
1,500 words (a long article) About 2,000 tokens
75,000 words (a full-length novel) About 100,000 tokens
750,000 words About 1 million tokens

Note: OpenAI says other languages “can have different relationships between characters, words, and tokens”. Text in Hindi, Tamil, Japanese or other non-English languages may need noticeably more tokens for the same meaning. Test your own text in a tokenizer before you budget.

Bar chart of output token prices per million: GPT-6 Astra $50, Claude Opus 5.5 $20, GPT-6.1 Sol $10, Claude Sonnet 5.5 $10, Claude Haiku 4.5 $5, Gemini 3.8 Flash $3.75, GPT-6 Luna $0.50
Output price per 1M tokens by model

Why tokens matter even if you never touch an API

Most people meet tokens through a chat app, not code. ChatGPT, Claude and Gemini do not show you a per-token bill on subscription plans, but tokens still shape what you get in three ways.

1. Usage limits are really token budgets

Anthropic is open about this. Its help center says the Pro plan allows “around 45 messages every five hours” for short conversations, and that the real number depends on message length, attachments, conversation length and model. A 20-page PDF and a long chat history are just a lot of tokens, so they use up your allowance faster. Our guide to Claude usage limits explains how the five-hour window works, and the same logic applies to ChatGPT limits by plan.

2. The context window is measured in tokens

The context window is how much text a model can consider at once: your instructions, the conversation so far, any files, and its reply. On paid Claude plans the base chat context is 200K tokens, rising to 1M tokens on models such as Opus 5.5 and Sonnet 5.5. OpenAI’s GPT-6 models accept 1,050,000 tokens through the API. When a conversation outgrows the window, older parts get dropped or summarized. For the full picture, read our context window explainer.

3. Every message resends the conversation

This is the part most people miss. A chat model has no memory between turns on its own. Each time you send a message, the app sends the whole conversation (or a trimmed version of it) back to the model. Your tenth message costs far more tokens than your first, even if it is one line long.

Input tokens, output tokens and the other types

Once you look at a pricing page, you will see several kinds of tokens. Here is what each one means.

  • Input tokens: everything you send to the model. That includes the system prompt (hidden instructions), the conversation history, uploaded documents and your actual question.
  • Output tokens: everything the model writes back. These cost more because generating text takes more computing work than reading it.
  • Cached input tokens: input the provider has seen recently and stored. If the start of your prompt is identical to a recent request, those tokens are billed at a steep discount. Our prompt caching guide covers how to set this up.
  • Reasoning tokens: internal “thinking” that reasoning models do before the visible answer. OpenAI confirms these count toward output usage and billing, even though you do not see them.
  • Audio and image tokens: media is also converted to tokens. OpenAI’s gpt-realtime-2.1 charges $32 per million audio input tokens and $64 per million audio output tokens, far more than its text rates. Google’s Gemini text-to-speech counts audio at 25 tokens per second.

Watch out: A short visible answer from a reasoning model can hide many more billed tokens. When you run GPT-6 Astra at high or max reasoning effort, the hidden thinking is charged at its $50 per million output rate.

How much do tokens cost? Real prices per million

API providers quote prices per 1 million tokens (often written as “per 1M” or “per MTok”). Here are current Standard-tier rates for popular models as of 30 September 2026.

Model Input per 1M Cached input per 1M Output per 1M
OpenAI GPT-6 Astra $10.00 $1.00 $50.00
OpenAI GPT-6.1 Sol $2.00 $0.10 $10.00
OpenAI GPT-6 Luna $0.10 $0.01 $0.50
Claude Opus 5.5 $4.00 $0.20 $20.00
Claude Sonnet 5.5 $2.00 $0.20 $10.00
Claude Haiku 4.5 $1.00 $0.10 $5.00
Gemini 3.8 Flash (until 31 Dec 2026) $0.75 $0.075 $3.75
Gemini 3.1 Flash-Lite $0.25 Not listed $1.50

Claude’s cached price shown is the cache read rate. Anthropic also charges a cache write fee of 1.25x the input price for a 5-minute cache. For the full rate cards, see our breakdowns of OpenAI API pricing, Claude API pricing and Gemini API pricing, or check the official OpenAI pricing page and Anthropic’s pricing docs.

Worked examples: turning tokens into dollars

The formula is simple: (input tokens / 1,000,000 x input price) + (output tokens / 1,000,000 x output price). Here it is applied to everyday jobs.

Example 1: writing a 1,000-word blog draft

Say your prompt with a brief and outline is 500 words (about 667 tokens) and the model writes 1,000 words (about 1,333 tokens).

  • GPT-6.1 Sol: (0.000667 x $2) + (0.001333 x $10) = $0.0013 + $0.0133 = about $0.015 per draft.
  • Claude Opus 5.5: (0.000667 x $4) + (0.001333 x $20) = $0.0027 + $0.0267 = about $0.029 per draft.
  • GPT-6 Luna: (0.000667 x $0.10) + (0.001333 x $0.50) = about $0.0007 per draft.

Even the premium option costs about 3 cents per draft. Tokens only add up at scale, which is why APIs suit agencies running hundreds of drafts. Our guide on writing blog posts with AI covers the quality side.

Example 2: summarizing a 50-page report

A 50-page report of about 25,000 words is roughly 33,000 tokens. You ask for a 400-word summary (about 533 tokens) using Claude Sonnet 5.5.

  • Input: 0.033 x $2 = $0.066
  • Output: 0.000533 x $10 = $0.0053
  • Total: about $0.071

Notice that input dominates here. When you send big documents, the input side of the bill matters most.

Example 3: a support chatbot at 1,000 chats a day

Each chat sends 2,000 input tokens and gets 500 output tokens back. That is 2M input and 0.5M output tokens per day.

  • GPT-6 Astra: (2 x $10) + (0.5 x $50) = $45 per day
  • GPT-6.1 Sol: (2 x $2) + (0.5 x $10) = $9 per day
  • GPT-6 Luna: (2 x $0.10) + (0.5 x $0.50) = $0.45 per day

Same traffic, a 100x difference in cost. Model choice is the biggest lever you have, which is why we wrote a full guide on picking the cheapest AI model for each task.

Save money: OpenAI, Anthropic and Google all offer a Batch API at 50% off for jobs that can wait up to 24 hours. The Sonnet 5.5 report job above would drop from $0.071 to about $0.036 on Claude’s batch rates. See our Batch API guide.

How to count tokens yourself

You do not have to guess. Each major provider gives you a way to count tokens before you send anything.

  1. Use a web tokenizer. OpenAI’s Tokenizer tool shows exactly how a piece of text splits into tokens, with each token highlighted in a different color. Paste in a typical prompt to see its real size.
  2. Use a library in code. OpenAI publishes tiktoken, an open-source library that counts tokens for its model encodings.
  3. Use Anthropic’s free counting endpoint. Claude’s API includes a token counting call that is free to use (with its own rate limits). It returns an estimate, which can differ slightly from the final count.
  4. Read the usage data. Every API response reports the input and output tokens it used, and each provider’s dashboard totals them. Check these numbers in your first week of real traffic.

Example: counting Claude tokens in Python

This uses Anthropic’s official SDK. The response is a small JSON object with the input token count.

import anthropic

client = anthropic.Anthropic()

response = client.messages.count_tokens(
    model="claude-opus-5-5",
    system="You are a helpful marketing assistant.",
    messages=[{"role": "user", "content": "Write 5 subject lines for a Diwali sale."}],
)

print(response.json())  # e.g. {"input_tokens": 31}

Example: a simple cost estimator

Once you know your token counts, a few lines of code turn them into money. Prices are per 1 million tokens.

PRICES = {
    "gpt-6.1-sol":       {"input": 2.00, "output": 10.00},
    "claude-sonnet-5-5": {"input": 2.00, "output": 10.00},
    "gpt-6-luna":        {"input": 0.10, "output": 0.50},
}

def cost(model, input_tokens, output_tokens):
    p = PRICES[model]
    return (input_tokens / 1_000_000) * p["input"] + (output_tokens / 1_000_000) * p["output"]

daily = cost(  # 1,000 chats a day, 2,000 tokens in and 500 out per chat
    "gpt-6.1-sol", 2_000 * 1_000, 500 * 1_000)
print(f"${daily:.2f} per day, ${daily * 30:.2f} per month")  # $9.00 per day, $270.00 per month

For a full planning method, including traffic assumptions and safety margins, read how to estimate AI API costs before you build.

Common myths about AI tokens

“One token equals one word”

Not quite. In English a token is about three-quarters of a word on average, so 1,000 words is closer to 1,333 tokens. Code, numbers, URLs and non-English text can use many more.

“Only my question counts”

Every request includes the system prompt, earlier messages and any attached files. In a long chat, your new question may be a tiny fraction of the tokens sent.

“A bigger context window is always better”

A larger window lets you send more, but you pay for everything you send. On OpenAI’s GPT-6 and GPT-5.6 models, requests over 272K input tokens are billed at 2x the input rate and 1.5x the output rate for the whole request. Sending only the relevant parts is cheaper and often gives better answers.

“Short answers are always cheap”

With reasoning models, the hidden thinking tokens can far exceed the visible reply. Lower the reasoning effort for simple tasks.

Practical ways to use fewer tokens

Whether you pay per token or work inside a subscription limit, the same habits help. We cover them in depth in our guide to reducing token usage in ChatGPT and Claude, but here are the essentials.

  • Start a new chat for a new topic. This stops old messages being resent every turn.
  • Upload only the pages you need. Paste the relevant section, not the whole PDF.
  • Ask for the format you want. “Five bullets, under 100 words” uses fewer output tokens than an open request.
  • Put reusable instructions in a project or system prompt. Anthropic says documents uploaded to a Claude Project are cached and count less against your limits than new content.
  • Write clearer prompts. A precise first prompt avoids three rounds of corrections. Our prompt engineering guide has 12 techniques.

Frequently asked questions

How many words is 1,000 tokens?

About 750 words of English text, using OpenAI’s rule that one token is roughly three-quarters of a word. The real number varies with word length, punctuation and formatting. Code, tables and non-English languages usually need more tokens per word, so check a sample of your own text in a tokenizer if you need an accurate count.

Why do output tokens cost more than input tokens?

Generating text is more work for the model than reading it, because each output token is produced one at a time. On current models output is typically five times the input price: GPT-6.1 Sol charges $2 input and $10 output per million tokens, and Claude Opus 5.5 charges $4 and $20. Asking for shorter answers is one of the easiest savings.

Do tokens matter on ChatGPT Plus or Claude Pro?

You are not billed per token on these subscriptions, but tokens still drive your usage limit. Anthropic says the number of Claude messages you get depends on message length, attachments, conversation length and model. Long chats and big files use more tokens per message, so you hit the cap sooner. The same applies to ChatGPT’s message caps.

Are tokens the same across ChatGPT, Claude and Gemini?

No. Each company uses its own tokenizer, so the same paragraph can produce a slightly different token count on each model. The rough English rule of about four characters per token is a fair estimate for all of them, but use each provider’s own counting tool, such as Anthropic’s free count endpoint, when precision matters.

What is the cheapest model per token right now?

Among the major labs, OpenAI’s GPT-6 Luna is one of the cheapest at $0.10 per million input tokens and $0.50 per million output tokens. Google’s Gemini 3.1 Flash-Lite costs $0.25 input and $1.50 output. DeepSeek’s API is also very cheap, and our DeepSeek vs ChatGPT comparison covers its trade-offs.

Next steps: put tokens to work

Tokens are just the unit AI uses to measure text. Once you know that 1,000 words is about 1,333 tokens, that output costs more than input, and that every chat message resends the history, both subscription limits and API bills start to make sense. The next step is choosing the right model for each job and trimming what you send.

Token prices change often, so always confirm current rates on each provider’s official pricing page before you budget.

Pricing and features are checked at the time of writing and can change. Some links may be affiliate links, which never affect our verdicts.

Written by

Ketan Parmar

Ketan Parmar has spent more than 15 years in digital marketing, helping brands grow through SEO, Google Ads, Meta Ads, content strategy and social media. Today he focuses on AI search visibility: how businesses get found and recommended in ChatGPT, Gemini, Perplexity and Google's AI answers.

Get the weekly AI tools brief

New tools, price changes and money-saving deals. One email a week, no spam.