• Monthly calculator
  • V4 Flash + Pro
  • Rates checked 18 Aug 2026

DeepSeek PricingTurn token rates into a monthly bill

DeepSeek pricing depends on model, input cache hits, output tokens and time of day. V4 Flash starts at $0.007 per million cached input tokens off-peak; V4 Pro starts at $0.022. Use the calculator below for a realistic monthly estimate.

Independent calculator · Official rates checked 18 August 2026 · Prices can change.

DeepSeek introduced the current V4 peak/off-peak table on 16 August 2026. Articles using the earlier flat prices do not calculate today's V4 bill correctly.

Estimate covers DeepSeek API tokens only. It excludes taxes, your infrastructure and third-party agent or tool charges.

DeepSeek V4 pricing at a glance

Flash cached input
$0.007 / 1M off-peak
Flash uncached input
$0.22 / 1M off-peak
Flash output
$0.66 / 1M off-peak
Pro cached input
$0.022 / 1M off-peak
Pro uncached input
$0.66 / 1M off-peak
Pro output
$1.98 / 1M off-peak
Peak premium
2× every token category
Context window
1M tokens

Interactive calculator

DeepSeek pricing calculator for monthly API usage

Enter average tokens per request and expected monthly requests. Cache-hit and peak-traffic percentages let the result reflect how the API is actually billed.

V4 Flash / month

$112.19$0.0112 per request

V4 Pro / month

$336.88$0.0337 per request

Flash savings vs Pro

$224.69for this monthly workload

Flash peak premium

$22.44extra cost caused by peak traffic

Estimate covers DeepSeek API tokens only. It excludes taxes, your infrastructure and third-party agent or tool charges.

Current API table

DeepSeek pricing per 1 million tokens

These are the V4 Flash and V4 Pro rates in DeepSeek's official documentation on 18 August 2026. Peak pricing is double the off-peak rate in every token category.

USD per 1 million tokens
Model and periodCached inputUncached inputOutput
V4 Flash · Off-peak$0.007$0.22$0.66
V4 Flash · Peak$0.014$0.44$1.32
V4 Pro · Off-peak$0.022$0.66$1.98
V4 Pro · Peak$0.044$1.32$3.96

V4 Flash

2,500 concurrent requests

Official concurrency limit shown for the Flash route.

V4 Pro

500 concurrent requests

Official concurrency limit shown for the Pro route.

Context

1M tokens

Both routes list a one-million-token context window.

Open official pricing

Time-of-day pricing

When DeepSeek peak pricing applies

Peak windows are 01:00–04:00 UTC and 06:00–10:00 UTC. Requests outside those seven hours use the off-peak table.

Orange: peakLight: off-peakTimeline: 00:00–24:00 UTC

Converted automatically

Your local peak windows are loading…

Actionable savings

How to lower DeepSeek pricing without cutting output quality

The rate table exposes two high-leverage levers: schedule flexible work off-peak and improve cache reuse on repeated context.

Move batch work outside peak windows

Every token category costs half as much off-peak. Queue evaluations, indexing and long reports for the lower-rate hours.

Keep repeated prefixes stable

System prompts, tool descriptions and long shared context are more likely to benefit from caching when their prefix does not change.

Route routine work to Flash

Use V4 Pro only when your own task tests show a quality gain worth roughly three times the off-peak uncached-input and output rates.

Measure cache hit rate

A long prompt is not automatically cheap. Read the provider usage breakdown, then enter the observed hit rate in the calculator.

FAQ

DeepSeek pricing questions

DeepSeek pricing questions from the reference page.

  • API tokens only
  • Peak is 2×
  • Verify current rates

There is no single monthly API fee in the official table. Monthly DeepSeek pricing is requests multiplied by cached input, uncached input and output charges, weighted by peak-hour traffic. Use the calculator above.

Sources

Official pricing source and limits

Independent calculator · Official rates checked 18 August 2026 · Prices can change.

Next step

Continue with DeepSeek

Practical pages for choosing and using agent infrastructure.