Blog / Model guides

DeepSeek V4 Flash: GPU requirements, boot time and cost

Published · Updated

Our DeepSeek V4 Flash deployment took just over 21 minutes to become ready on two H200 GPUs. That is a useful starting point for planning a private endpoint, especially if you intend to shut it down between uses. It is a measurement of one configuration; the model revision, download path and available capacity all affect the next boot.

This guide retains our August 2026 deployment observations. DeepSeek has since published updates, including the 0731 build and its DSpark speculative-decoding module. Check the official changelog and the exact checkpoint before applying these figures to a new deployment.

The hardware we used

ItemMeasured configuration
GPUs2x NVIDIA H200, 141 GB per GPU
Provider and locationRunPod, Iceland, in our own account
Weights downloaded159.6 GB in 46 shards, with mixed FP8 and FP4 expert weights
Served context32,768 tokens
Cold boot on 11 August 202621 minutes 26 seconds
Repeat boot on 18 August 2026About 21 minutes

Iceland is in the EEA, but it is not an EU member state. Choose an EU region if your deployment policy requires one.

The measured checkpoint is too large for one 141 GB H200 before adding the memory needed to serve requests. The catalog also lists a 4x H100 80 GB alternative. GPU memory is only part of the fit check: the engine must support the checkpoint's quantization on that GPU. Our qualified configuration uses Hopper hardware; do not assume an A100 can run the same image.

Keep the model's advertised context separate from the context you actually serve. Longer conversations and more simultaneous requests need extra cache memory. A deployment that loads successfully can still run out of memory under its intended workload.

What the boot measurement tells you

The clock runs from creating a fresh instance to a ready endpoint. It includes downloading weights, loading the engine and preparing it to serve. It does not measure how quickly the model answers a busy application.

We have not published a workload-controlled throughput result for these runs. For a useful comparison, record input length, output length, concurrency, time to first token and completed requests. Test a representative long request as well as the short chat you use to check that the endpoint works.

Plan for the whole bill

We observed RunPod rates of $7 to $14 per hour for the two-GPU shape during the August boots. At those historical rates, the compute arithmetic is:

Usage assumptionCompute cost
720 hours, a 30-day month$5,040 to $10,080
176 serving hours$1,232 to $2,464, before startup time
One 21-minute-26-second startupAbout $2.50 to $5.00, if billed throughout

Add retained storage, any network charges and the LLM Hangar subscription. Your provider sets the actual rate and bills you directly. The cost calculator lets you change the rate, schedule and startup allowance. Budget caps and timers help limit spending, but they use estimates and do not guarantee an exact invoice.

When does a private endpoint make financial sense?

Use your actual API bill as the comparison. Separate input, output and cached tokens, then include any time-of-day pricing. DeepSeek's pricing announcements make this more useful than a single blended rate copied from an older guide.

For an illustrative calculation, suppose your avoided API cost is $1 per million tokens and your GPU node costs $10.50 an hour. You would need to replace 10.5 million billable tokens each running hour to cover compute alone. That is roughly 2,917 combined input and output tokens per second. It is a required rate, not a measured capability of this model.

Then ask whether the endpoint can sustain your request mix at an acceptable latency. A favorable token calculation means little if queues grow during the hours people need it. Low or irregular usage may suit an API better; a private endpoint may still be worth the cost for control over the request path and deployment configuration.

Current prices, live

Loading current prices…

Deploy and check the client connection

  1. Connect your cloud account and check that your plan and provider support the required GPU shape.
  2. Select the model, region and any fallback shapes. Review the estimated rate and set a budget cap.
  3. Choose staging if you expect to stop and resume regularly. Include its storage charge.
  4. Once ready, use the endpoint URL and key from the Connect tab. Test streaming, long requests and any tool calls your application needs.

The endpoint supports compatible chat-completion clients. See client examples for setup. With the default node-local gateway, requests go directly to your instance. A hosted edge changes that path; review gateway placement and access before sending sensitive data.

Frequently asked questions

Can the measured checkpoint run on one H200?

No. The 159.6 GB checkpoint we downloaded exceeds one H200's 141 GB memory before runtime overhead. We measured it on two H200s.

Does the 21-minute boot time apply to the latest build?

It applies to our catalog configuration on 11 and 18 August 2026. A different checkpoint, engine, provider or download path needs its own measurement.

Does a schedule remove all costs outside working hours?

Stopping removes GPU charges once the provider confirms the compute is gone. Retained storage, an optional private edge and the platform subscription can still cost money. Resuming also takes time and depends on capacity.