LLM Hangar / Tools / LLM hosting cost calculator

LLM hosting cost calculator: GPU cost per model

Updated .

Choose a model to estimate a GPU configuration and monthly cost. Compare schedules, startup time and retained-weight storage using the available price data. Memory fit is an estimate; check engine compatibility before renting hardware.

01

Pick a model

Size
Weights1
Context2

1 As shipped is the format the weights are published in: fp8 for DeepSeek V4 Flash and GLM-5.3-Flash, MXFP4 for Kimi K3 and gpt-oss, bf16 for most others. Lower precision usually uses less memory; its effect on answer quality depends on the model, conversion and task. 2 The longest request the endpoint must hold in cache.

02

Minimum required hardware

Hardware
Memory needed
Per hour
03

Cost of private DeepSeek V4 Flash

Schedule3
Line
Quantity
Per month
GPU hours
of which startup time, before the first token on each start
Staged weights
LLM Hangar planflat, does not change with hours; see pricing
1 plan
separate
Provider billGPU hours plus storage, at the rates above

Compare this budget with an API and add operating costs

Prices come from the LLM Hangar price index, last checked 28 August 2026, 04:00 UTC. For a model the index lists, the hardware is the cheapest of its checked offers and provider estimates that holds the memory needed; for the others it is the cheapest per-GPU rate for each card type in the same index. Your provider bills you at its own rates and none of this is a quotation.

3 Both cost calculators use a 730-hour planning month. Eight-hour weekdays use 176 serving hours and 22 starts; ten-hour business days (08:00 to 18:00) use about 217 hours and 22 starts. Weekdays 24x5 use about 522 hours and 5 weekly starts. Always-on startup is included in 730 billed hours. Other presets add startup before serving; adjust Custom for your calendar or first-month downtime. Staged weights can shorten startup, so enter your measured warm boot. Storage is estimated as max(weights x 1.1, 100 GB) at the provider's list rate per GB-month; a confirmed deployment shows its actual stage quote.

memory = weights + KV cache, plus 15% headroom weights = parameters x bytes per parameter (4-bit 0.55, 8-bit 1.05, 16-bit 2.0) at the shipped or chosen precision, or the published checkpoint size where we have one KV cache = 50 KB x sqrt(active parameters in billions) x context tokens, or the model's own bytes per token where we have read its config hardware = cheapest card type and count (1, 2, 4, 8, 16) whose memory covers it GPU hours = serving hours + starts x boot minutes / 60 (always on 730 including startup, 8-hour weekdays 176, business hours 217, weekdays 24x5 522) storage = max(weights x 1.1, 100 GB) x $/GB-month, shown whenever the weights are staged provider bill = $/h x GPU hours + storage; the plan fee is flat and separate
Where the numbers come from
  • Model sizes and licences: the model cards on Hugging Face and vendor documentation for DeepSeek V4 Flash and V4 Pro, GLM-5.3-Flash, GLM-5.2, Kimi K3, Qwen3.8-2.4T-A95B, Qwen3.8-27B, the Qwen3.5 family, the Gemma 4 model card, Muse Glimmer 30B, Mistral Small 4, gpt-oss, Llama 4. Sizes for Mistral Large 3, Kimi K2.5 and Qwen3 235B are as listed by Thunder Compute on 20 August 2026. The calculator still uses GLM-5.2-based sizing for GLM-5.3. The flagship now has published weights; use its GLM-5.3: hardware requirements and license and exact checkpoint for a deployment decision.
  • Checkpoint sizes: DeepSeek V4 Flash 159.6 GB fp8, GLM-5.3-Flash 328 GB fp8 and Gemma 4 31B 58.25 GiB bf16 as downloaded in the measured deployments above; Kimi K3 1,560.94 GB MXFP4 and Qwen3.8-2.4T-A95B 4.89 TB bf16 from the shard listings on Hugging Face.
  • Prices: the LLM Hangar price index, which records checked offers and provider list prices across AWS, Azure, GCP, OCI, Nebius, RunPod, Lambda, Verda and CloudRift every eight hours. The page loads the latest check and falls back to the 28 August 2026 snapshot.
  • Boot times: measured cold boots on RunPod, from instance creation to a ready endpoint, 21 minutes 26 seconds for DeepSeek V4 Flash: GPU requirements, boot time and cost (11 August 2026), 34 minutes for GLM-5.3-Flash (27 August 2026) and about 14 minutes for Gemma 4 31B on one H100 (24 August 2026). Those models prefill their own figure; the 20-minute default for the others is an assumption inside that range, and the field is editable.

Questions

How much GPU memory does a model need?

Weights plus KV cache plus headroom. Weights are the parameter count times the bytes per parameter: about 2 bytes at 16-bit, 1 at 8-bit and 0.55 at 4-bit including the quantisation tables. The KV cache grows with context length and with the model's attention layout. Measured examples: Gemma 4 31B in bf16 loads 57.91 GiB of weights and needs 15.79 GiB of KV cache for one 32,768-token request; the DeepSeek V4 Flash fp8 checkpoint is 159.6 GB and runs on 2x H200.

How much does it cost to host an LLM per month?

The hourly rate of the cheapest GPU configuration in the price index that holds the model, times the hours it exists. Always on is 730 hours a month; 8 hours a day on weekdays is 176 plus a cold start each working day. At the index check of 28 August 2026, DeepSeek V4 Flash on 2x H200 at $8.00 an hour is $5,840 a month always on and about $1,470 for eight-hour weekdays including 22 boots; Kimi K3 at its native MXFP4 needs 16x H200 at $64 an hour, about $46,700 always on. The GPU bill comes from your provider; any platform fee is on top.

What does a cold boot cost?

The GPU meter runs from instance creation, before the first token. Measured cold boots on RunPod: 21 minutes 26 seconds for DeepSeek V4 Flash on 2x H200, 34 minutes for GLM-5.3-Flash on 4x H200 and about 14 minutes for Gemma 4 31B on one H100. At the $8.00 an hour the price index listed for 2x H200 on 28 August 2026, a 21-minute boot is $2.80 of dead time per start, about $62 a month if the box is started cold every working day. A stopped instance still bills its disk.

Start a 7-day free trial