Blog / Model guides
DeepSeek V4 Flash: GPU requirements, boot time and cost
Our DeepSeek V4 Flash deployment took just over 21 minutes to become ready on two H200 GPUs. That is a useful starting point for planning a private endpoint, especially if you intend to shut it down between uses. It is a measurement of one configuration; the model revision, download path and available capacity all affect the next boot.
This guide retains our August 2026 deployment observations. DeepSeek has since published updates, including the 0731 build and its DSpark speculative-decoding module. Check the official changelog and the exact checkpoint before applying these figures to a new deployment.
The hardware we used
| Item | Measured configuration |
|---|---|
| GPUs | 2x NVIDIA H200, 141 GB per GPU |
| Provider and location | RunPod, Iceland, in our own account |
| Weights downloaded | 159.6 GB in 46 shards, with mixed FP8 and FP4 expert weights |
| Served context | 32,768 tokens |
| Cold boot on 11 August 2026 | 21 minutes 26 seconds |
| Repeat boot on 18 August 2026 | About 21 minutes |
Iceland is in the EEA, but it is not an EU member state. Choose an EU region if your deployment policy requires one.
The measured checkpoint is too large for one 141 GB H200 before adding the memory needed to serve requests. The catalog also lists a 4x H100 80 GB alternative. GPU memory is only part of the fit check: the engine must support the checkpoint's quantization on that GPU. Our qualified configuration uses Hopper hardware; do not assume an A100 can run the same image.
Keep the model's advertised context separate from the context you actually serve. Longer conversations and more simultaneous requests need extra cache memory. A deployment that loads successfully can still run out of memory under its intended workload.
What the boot measurement tells you
The clock runs from creating a fresh instance to a ready endpoint. It includes downloading weights, loading the engine and preparing it to serve. It does not measure how quickly the model answers a busy application.
We have not published a workload-controlled throughput result for these runs. For a useful comparison, record input length, output length, concurrency, time to first token and completed requests. Test a representative long request as well as the short chat you use to check that the endpoint works.
Plan for the whole bill
We observed RunPod rates of $7 to $14 per hour for the two-GPU shape during the August boots. At those historical rates, the compute arithmetic is:
| Usage assumption | Compute cost |
|---|---|
| 720 hours, a 30-day month | $5,040 to $10,080 |
| 176 serving hours | $1,232 to $2,464, before startup time |
| One 21-minute-26-second startup | About $2.50 to $5.00, if billed throughout |
Add retained storage, any network charges and the LLM Hangar subscription. Your provider sets the actual rate and bills you directly. The cost calculator lets you change the rate, schedule and startup allowance. Budget caps and timers help limit spending, but they use estimates and do not guarantee an exact invoice.
When does a private endpoint make financial sense?
Use your actual API bill as the comparison. Separate input, output and cached tokens, then include any time-of-day pricing. DeepSeek's pricing announcements make this more useful than a single blended rate copied from an older guide.
For an illustrative calculation, suppose your avoided API cost is $1 per million tokens and your GPU node costs $10.50 an hour. You would need to replace 10.5 million billable tokens each running hour to cover compute alone. That is roughly 2,917 combined input and output tokens per second. It is a required rate, not a measured capability of this model.
Then ask whether the endpoint can sustain your request mix at an acceptable latency. A favorable token calculation means little if queues grow during the hours people need it. Low or irregular usage may suit an API better; a private endpoint may still be worth the cost for control over the request path and deployment configuration.
Current prices, live
Deploy and check the client connection
- Connect your cloud account and check that your plan and provider support the required GPU shape.
- Select the model, region and any fallback shapes. Review the estimated rate and set a budget cap.
- Choose staging if you expect to stop and resume regularly. Include its storage charge.
- Once ready, use the endpoint URL and key from the Connect tab. Test streaming, long requests and any tool calls your application needs.
The endpoint supports compatible chat-completion clients. See client examples for setup. With the default node-local gateway, requests go directly to your instance. A hosted edge changes that path; review gateway placement and access before sending sensitive data.
Frequently asked questions
Can the measured checkpoint run on one H200?
No. The 159.6 GB checkpoint we downloaded exceeds one H200's 141 GB memory before runtime overhead. We measured it on two H200s.
Does the 21-minute boot time apply to the latest build?
It applies to our catalog configuration on 11 and 18 August 2026. A different checkpoint, engine, provider or download path needs its own measurement.
Does a schedule remove all costs outside working hours?
Stopping removes GPU charges once the provider confirms the compute is gone. Retained storage, an optional private edge and the platform subscription can still cost money. Resuming also takes time and depends on capacity.