Compare the full monthly cost of serving acceptable answers. This guide uses official API rates checked October 4, 2026 and clearly labeled illustrative calculations. We did not conduct the previously claimed 30-day, 500-query-per-day experiment.

Choose with your workload: A low-priced API can remain cheaper even at substantial token volume. Self-hosting becomes attractive when utilization, quality, latency, control and operating costs work together. There is no universal 5M, 10M or 100M-token break-even point.

Start with three separate token counters

A request count is not a bill. Record uncached input tokens, cached input tokens and output tokens separately. Include reasoning/output billing according to the selected provider's documentation, retries and unsuccessful jobs. A long-input summarizer and a short-input code generator can have very different bills at the same request volume.

API monthly cost = (uncached input × input rate + cached input × cache rate + output × output rate) / 1,000,000 + charged tools, storage and other services.

Use the actual billed cache-hit counter. Similar prompts do not guarantee cache hits. Count retries in the token totals rather than adding their cost twice. If traffic crosses tariff periods, compute each period separately.

Current API example: DeepSeek Flash and Pro

DeepSeek's current catalog maps deepseek-flash to DeepSeek-V4.1-Flash and retains deepseek-v4-pro for DeepSeek-V4-Pro-0813. An earlier retirement plan was corrected in the current changelog. API availability does not establish which model a consumer chat account uses.

USD per one million tokens; checked October 4, 2026
API modelPeriodCached inputUncached inputOutput
deepseek-flash (V4.1 Flash)Off-peak$0.003$0.15$0.60
deepseek-flash (V4.1 Flash)Peak$0.006$0.30$1.20
deepseek-v4-pro (V4-Pro-0813)Off-peak$0.022$0.66$1.98
deepseek-v4-pro (V4-Pro-0813)Peak$0.044$1.32$3.96

Peak periods: 01:00–04:00 and 06:00–10:00 UTC on Monday–Friday, excluding Chinese public holidays. The remaining hours, weekends and those holidays are off-peak under the current rules. These are API rates, separate from chat subscriptions. Model features also differ: Flash supports image inputs; Pro's current row does not.

A reproducible monthly API calculation

Illustrative workload, not measured usage: 8 million cached input tokens, 2 million uncached input tokens and 1 million output tokens, all sent to Flash in the same tariff period. This assumes that the usage report actually records those 8 million cache hits.

Off-peak: 8 × $0.003 + 2 × $0.15 + 1 × $0.60 = $0.924.
Peak: 8 × $0.006 + 2 × $0.30 + 1 × $1.20 = $1.848.

These totals cover only the stated model tokens. They exclude taxes, tools, additional retries and other infrastructure. Without any cache hits, the same 10 million input tokens plus 1 million output would cost $2.10 off-peak or $4.20 peak. Replace every assumption with your billing counters before making a purchase decision.

What belongs in a local model budget?

  • Compute: Billed GPU hours at your actual contract rate, including idle time and redundancy. For owned hardware, use a reasonable monthly allocation of purchase cost, power and maintenance.
  • Supporting infrastructure: CPU, RAM, storage, network transfer, monitoring and backups.
  • Operations: Engineering time for deployment, evaluation, upgrades, capacity management and incident recovery, valued at your own staffing cost.
  • Fallback: Any API requests used when the local model fails, overloads or cannot meet the quality target.

Local monthly cost = billed compute + supporting infrastructure + operating labor + fallback usage.

Idle time is already included in billed instance hours; avoid adding an arbitrary idle markup on top. There is no verified universal $2/hour GPU rate or fixed engineering multiplier in this guide. Obtain a dated quote for the region, accelerator, memory, availability commitment and minimum billing duration you will actually use.

Compare cost per accepted task

Token savings are useful only if the output meets the job's acceptance criteria. Use the same task set and human or automated checks for both deployments. Measure completed, acceptable tasks rather than counting every generated response as success.

Effective cost per accepted task = total serving and review cost / number of accepted tasks.

Record correctness, instruction compliance, p95 completion latency, peak concurrency and retry rate. Quantization, context length and batching affect both memory and quality. A GPU memory figure alone does not prove that a given model can serve your required context and concurrency. Benchmark your proposed model, configuration and hardware combination.

Calculate a workload-specific break-even point

If a local deployment has fixed monthly cost F, an API costs A per accepted task and local variable cost is V per accepted task, then:

Break-even accepted tasks = F / (A − V), provided A > V.

This simplified equation assumes comparable quality and enough capacity. If A ≤ V, increasing traffic will not recover the fixed cost under those assumptions. Add another GPU, staffing change or latency constraint and the equation must be recalculated. Do not translate it into a universal monthly token threshold.

A practical pilot before migration

  1. Choose a representative sample of real tasks and define acceptance criteria before comparing outputs.
  2. Measure input/cache/output tokens and chargeable requests over a representative period, including a busy interval.
  3. Test the proposed local configuration under concurrent traffic. Record quality, queueing, failures and recovery.
  4. Apply actual dated API tariffs and compute quotes, then include labor and fallback costs.
  5. Compare monthly totals and cost per accepted task. Document how results change when demand doubles or utilization drops.

Use hybrid routing when it earns its complexity

A local model may handle repetitive extraction or classification while an API handles cases it cannot complete satisfactorily. Choose the split using observed acceptance rates and routing cost. No fixed 70/30 split or guaranteed 60–70% saving is established here.

Measure false routing, duplicate inference, fallback usage and operational effort. A simpler API deployment can be preferable when these costs exceed the benefit. A local runtime such as Ollama can help run a pilot, but installation convenience does not establish production capacity or security.

Frequently asked questions

Is local always cheaper above 10 million tokens?

No. Input/output mix, cache hits, tariff periods, model quality, utilization and staffing determine the result. Use current rates and accepted-task costs.

Does self-hosting guarantee private data?

No. Review access controls, logs, backups, integrations and network paths. Provider terms and your deployment design require separate checks; a generic privacy score cannot replace them.

Should I buy a GPU before measuring usage?

First measure workload and test a compatible configuration. Hardware memory requirements and concurrency must be verified for the model, precision and context you intend to serve.

Sources and scope

Editorial analysis based on primary documentation; the numerical workload is illustrative. Rates were checked October 4, 2026 and can change. This is a corrected cost framework, not a record of hands-on benchmarking.

Read the DeepSeek profile for availability caveats, or browse AIListPrime's full tools list →.