Skip to content
Pricing and Plans

DeepSeek API Pricing 2026: Cost Per Million Tokens Explained

DeepSeek API pricing for 2026: deepseek-flash and V4-Pro costs per million tokens, peak and off-peak hours in IST, cache hits, worked examples and how it compares.

DeepSeek API Pricing 2026: Cost Per Million Tokens Explained
On this page
  1. Key takeaways
  2. DeepSeek API pricing table
  3. Peak and off-peak hours explained
  4. deepseek-flash vs deepseek-v4-pro: which model to use
  5. How cache hits work on DeepSeek
  6. Worked cost examples
  7. DeepSeek vs OpenAI, Gemini and Claude API pricing
  8. Hidden costs and limits to plan for
  9. How to pay less for the DeepSeek API
  10. Is DeepSeek safe for business data?
  11. Which DeepSeek option should you pick?
  12. Frequently asked questions
  13. Verdict: cheap, flexible, but compare before you commit

The DeepSeek API costs $0.15 per million input tokens and $0.60 per million output tokens for its main model, deepseek-flash (DeepSeek-V4.1-Flash), at off-peak times, and double that ($0.30 and $1.20) during peak hours. The larger deepseek-v4-pro costs $0.66 input and $1.98 output off-peak, or $1.32 and $3.96 at peak. Cached input is far cheaper still, from $0.003 per million tokens.

This guide covers every DeepSeek API price as of October 2026, when peak hours fall (including Indian time), how cache hits work, worked cost examples, how DeepSeek compares with OpenAI, Google and Anthropic, and the hidden costs to plan for.

Key takeaways

  • Two models: deepseek-flash (V4.1-Flash) and deepseek-v4-pro, both with a 1M token context window and up to 384K output tokens.
  • Off-peak rates are half of peak. Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, which is 6:30 to 9:30 am and 11:30 am to 3:30 pm in India.
  • Cache hits cost 2% of the cache-miss price on Flash ($0.003 vs $0.15 off-peak), so repeated prompt prefixes save a lot.
  • V4-Pro is still available at its listed prices: DeepSeek reversed an earlier plan to route all V4-Pro requests to Flash.
  • GPT-6 Luna ($0.10 input, $0.50 output) undercuts DeepSeek Flash at peak, so DeepSeek is not automatically the cheapest option.
Bar chart of DeepSeek output prices per million tokens: Flash off-peak $0.60, Flash peak $1.20, V4-Pro off-peak $1.98, V4-Pro peak $3.96
DeepSeek output price per 1M tokens

DeepSeek API pricing table

All prices are USD per 1 million tokens, from the official DeepSeek models and pricing page.

Model Time Input (cache hit) Input (cache miss) Output
deepseek-flash (V4.1-Flash) Off-peak $0.003 $0.15 $0.60
deepseek-flash (V4.1-Flash) Peak $0.006 $0.30 $1.20
deepseek-v4-pro (V4-Pro-0813) Off-peak $0.022 $0.66 $1.98
deepseek-v4-pro (V4-Pro-0813) Peak $0.044 $1.32 $3.96

Both models support a 1M token context window and a maximum output of 384K tokens, plus thinking mode (the default) and non-thinking mode. If tokens are new to you, our plain English guide to AI tokens explains them; roughly, 1 million tokens is about 750,000 English words.

Peak and off-peak hours explained

DeepSeek prices by time of day. Peak hours are 01:00 to 04:00 UTC and 06:00 to 10:00 UTC, Monday to Friday, excluding Chinese public holidays. Every other hour, including weekends and Chinese public holidays, is off-peak at 50% of peak rates.

Window UTC India (IST) Price level
Peak 1 (weekdays) 01:00 to 04:00 6:30 am to 9:30 am Full price
Peak 2 (weekdays) 06:00 to 10:00 11:30 am to 3:30 pm Full price
All other weekday hours Rest of the day 9:30 to 11:30 am, and 3:30 pm to 6:30 am Half price
Weekends and Chinese holidays All day All day Half price

For Indian teams, the peak windows overlap the working morning and early afternoon. Anything that can run in the evening, overnight or on weekends (bulk content, data extraction, embeddings preparation, test runs) costs half as much.

Save money: Schedule bulk jobs after 3:30 pm IST on weekdays or at weekends. There is nothing to configure: DeepSeek applies the off-peak rate automatically based on when the request runs.

deepseek-flash vs deepseek-v4-pro: which model to use

deepseek-flash (DeepSeek-V4.1-Flash)

Flash is DeepSeek’s newest model, released on 10 September 2026, and the one most people should use. DeepSeek describes it as the smallest model in its new architecture family, with native visual understanding, and it supports image input, JSON output, tool calls, the Responses API format and an Anthropic-compatible API. DeepSeek claims it beats V4-Pro on its published benchmarks, and the open weights are on Hugging Face. Its concurrency limit is 2,500 simultaneous requests per account.

Use it for: chatbots, summarization, extraction, coding help, content drafts and most agent work. Skip it if: you need a model whose behavior you have already validated on V4-Pro and do not have time to retest.

deepseek-v4-pro (V4-Pro-0813)

V4-Pro reached general availability in August 2026. It costs about 4.4 times more than Flash per input token and 3.3 times more per output token, has no vision support, and a lower concurrency limit of 500. In September, DeepSeek announced that V4-Pro requests would be routed to V4.1-Flash from 14 September until a future V4.1-Pro launches, but its change log now says it “decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026” with billing unchanged.

Use it for: existing pipelines tuned on V4-Pro. Skip it if: you are starting fresh; test Flash first, since it is cheaper and newer.

Watch out: DeepSeek says V4-Pro is being phased out and that it will give notice of further changes. Old model names such as deepseek-v4-flash are retired and now served by V4.1-Flash, so update your code to use deepseek-flash.

How cache hits work on DeepSeek

DeepSeek charges a much lower “cache hit” rate when the start of your prompt matches a prompt it has recently processed, and the full “cache miss” rate for new input. On Flash, a cache hit costs $0.003 per million tokens off-peak versus $0.15 for a miss, a 98% discount.

To get more cache hits, keep everything that repeats (system prompt, instructions, examples, reference documents) at the very start of each request, and put the part that changes (the user’s question) at the end. The same principle applies across providers, as our prompt caching guide explains.

Worked cost examples

Example 1: writing 1,000 articles

Each article uses 1,500 input tokens (brief and instructions) and 2,500 output tokens. That is 1.5M input and 2.5M output, with no caching.

  • Flash, off-peak: (1.5 x $0.15) + (2.5 x $0.60) = $0.225 + $1.50 = about $1.73
  • Flash, peak: (1.5 x $0.30) + (2.5 x $1.20) = $0.45 + $3.00 = $3.45
  • V4-Pro, off-peak: (1.5 x $0.66) + (2.5 x $1.98) = $0.99 + $4.95 = $5.94
  • V4-Pro, peak: (1.5 x $1.32) + (2.5 x $3.96) = $1.98 + $9.90 = $11.88

Example 2: a support bot with caching

A support bot processes 50M input tokens and 5M output tokens a month, and 80% of the input (the long system prompt and help-center content) hits the cache.

  • Flash, peak, with cache: 40M hits x $0.006 = $0.24, plus 10M misses x $0.30 = $3.00, plus 5M output x $1.20 = $6.00. Total $9.24.
  • Flash, peak, no cache: 50M x $0.30 = $15.00, plus $6.00 output. Total $21.00.
  • Flash, off-peak, with cache: half of $9.24, so $4.62.

Caching more than halves the bill, and output becomes the biggest cost. Keeping answers concise is the next lever. To model your own numbers, follow our guide on estimating AI API costs before you build.

DeepSeek vs OpenAI, Gemini and Claude API pricing

DeepSeek is cheap, but no longer in a class of its own. Here is how its prices compare with popular models (standard rates per million tokens, cache-miss input).

Model Input Output Context
GPT-6 Luna (OpenAI) $0.10 $0.50 1.05M
DeepSeek Flash, off-peak $0.15 $0.60 1M
Gemini 3.1 Flash-Lite (text) $0.25 $1.50 See Google docs
DeepSeek Flash, peak $0.30 $1.20 1M
DeepSeek V4-Pro, off-peak $0.66 $1.98 1M
Claude Haiku 4.5 $1.00 $5.00 200K
DeepSeek V4-Pro, peak $1.32 $3.96 1M
GPT-6.1 Sol / Claude Sonnet 5.5 $2.00 $10.00 1.05M / 1M

Running Example 2 (50M input with 80% cache hits, 5M output) on GPT-6 Luna, with its $0.01 cached input rate, costs (40 x $0.01) + (10 x $0.10) + (5 x $0.50) = $0.40 + $1.00 + $2.50 = $3.90, slightly less than DeepSeek Flash off-peak. On the other hand, DeepSeek’s 384K maximum output is far larger than the 128K offered by GPT-6 models, which helps when generating very long files.

Prices only tell half the story; quality differs, so test your own prompts. Our guides to OpenAI API pricing, Gemini API pricing and Claude API pricing list full rate cards, and DeepSeek vs ChatGPT compares quality and privacy.

Hidden costs and limits to plan for

  • Peak pricing surprises. A job that usually runs overnight but slips into the morning in India costs double. Log request times if you budget tightly.
  • Thinking mode output. Thinking mode is the default on both models. Reasoning text generally adds to the output you pay for, and output is the most expensive token type, so test non-thinking mode for simple tasks such as classification.
  • Concurrency limits. Exceed 2,500 concurrent requests on Flash or 500 on V4-Pro and you get an HTTP 429 error. DeepSeek accepts capacity expansion requests at no extra cost. If a request has not started processing after 10 minutes, the server closes the connection.
  • Prepaid balance. Usage is deducted from a topped-up balance (any granted balance is used first), so set reminders to top up before production traffic stops.
  • Price changes. DeepSeek’s page states that prices “may vary” and that it reserves the right to adjust them. It cut prices with the V4.1-Flash launch, and could change them again.

Tip: DeepSeek’s Anthropic-compatible endpoint maps Claude model names to its own models (Opus names to v4-pro, Sonnet and Haiku names to flash). It is a quick way to test DeepSeek in Anthropic-format tools, but several features such as cache_control are ignored. Our guide to reducing Claude Code and Cursor costs covers the trade-offs.

How to pay less for the DeepSeek API

  1. Default to Flash. It is the cheapest and newest model, with vision and tool calls.
  2. Shift bulk work off-peak. Evenings, nights and weekends in India are half price.
  3. Structure prompts for cache hits. Fixed content first, changing content last.
  4. Cap output length. Ask for concise answers and set a sensible maximum, since output costs four times more than cache-miss input.
  5. Route by difficulty. Send easy requests to the cheapest capable model, whichever provider that is. Our explainer on model routing shows how to mix DeepSeek with other APIs.
  6. Consider self-hosting. V4.1-Flash weights are published, but a 552B-parameter model needs serious hardware. For local experiments, smaller open models are more practical; see open source AI models you can use for free.

Is DeepSeek safe for business data?

DeepSeek is a Chinese company, and some governments and organizations have reportedly restricted its apps on work devices. Before sending customer data, contracts or proprietary code, read its current privacy policy and terms, check where data is stored, and make sure your contracts and local laws allow it. We could not verify DeepSeek’s data storage details for this guide. Our AI data privacy guide for businesses lists the questions to ask any AI vendor.

Which DeepSeek option should you pick?

You are… Best choice Why
A student or hobbyist testing APIs Flash, off-peak Lowest cost, full features
A content team generating drafts in bulk Flash, scheduled off-peak About $1.73 per 1,000 articles in our example
A startup running a support bot Flash with cached prompts Cache hits cost 2% of fresh input
A team with V4-Pro pipelines V4-Pro, while testing Flash Still available, but being phased out
A business with sensitive data Review privacy first Consider US or EU providers or self-hosting

Frequently asked questions

How much does the DeepSeek API cost?

deepseek-flash costs $0.15 per million input tokens and $0.60 per million output tokens off-peak, or $0.30 and $1.20 at peak. deepseek-v4-pro costs $0.66 and $1.98 off-peak, or $1.32 and $3.96 at peak. Cache hits are much cheaper, from $0.003 per million tokens on Flash off-peak.

When are DeepSeek off-peak hours?

Off-peak is every hour outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, plus all weekends and Chinese public holidays. In India that means weekday peak runs 6:30 to 9:30 am and 11:30 am to 3:30 pm IST. Off-peak prices are half of peak and apply automatically.

Is the DeepSeek API cheaper than OpenAI?

Not always. GPT-6 Luna costs $0.10 input and $0.50 output, which is cheaper than DeepSeek Flash at peak and slightly cheaper off-peak. DeepSeek Flash is far cheaper than mid-tier models such as GPT-6.1 Sol or Claude Sonnet 5.5 at $2 and $10. Compare quality on your own tasks before deciding.

Is DeepSeek V4-Pro still available?

Yes. DeepSeek first said V4-Pro requests would be routed to V4.1-Flash from 14 September 2026, then updated its change log to say it would continue V4-Pro API service with billing unchanged. It also says V4-Pro is being phased out and will give notice of changes, so plan to test Flash.

Does DeepSeek have a free API tier?

DeepSeek’s pricing page mentions a “granted balance” that is used before your topped-up balance, but we could not verify a standard free allowance for new accounts. Expect to top up a prepaid balance in USD. The free DeepSeek chat app is a separate product from the paid API.

Verdict: cheap, flexible, but compare before you commit

DeepSeek’s API remains one of the cheapest ways to access a capable model with a 1M token context window and very long outputs. deepseek-flash at $0.15 input and $0.60 output off-peak, with cache hits from $0.003, suits bulk content, support bots and coding helpers, especially for Indian teams who can schedule work after 3:30 pm. It is no longer automatically the cheapest option, since GPT-6 Luna undercuts it, and privacy reviews matter for business data. Use Flash by default, keep V4-Pro only for existing pipelines, and check the official DeepSeek pricing page before you budget, because prices and models change often.

Related reading: the cheapest AI model for each task and how to run an LLM locally with Ollama.

Pricing and features are checked at the time of writing and can change. Some links may be affiliate links, which never affect our verdicts.

Written by

Ketan Parmar

Ketan Parmar has spent more than 15 years in digital marketing, helping brands grow through SEO, Google Ads, Meta Ads, content strategy and social media. Today he focuses on AI search visibility: how businesses get found and recommended in ChatGPT, Gemini, Perplexity and Google's AI answers.

Get the weekly AI tools brief

New tools, price changes and money-saving deals. One email a week, no spam.