Docs / Deploy a model
Deploy a model
Choose a model
The first step asks what you want to deploy. LLM is a language model behind an OpenAI-compatible API, and the rest of this page describes it. RLCD is a decision model, which answers typed questions with probabilities; see decision models.
The catalog lists curated open models (DeepSeek, GLM, Kimi, Qwen, Gemma and others) with a GPU shape that is known to run each one. You can also deploy a model from a Hugging Face repository: paste the repo id and pick a shape yourself.
New custom deployments use vLLM 0.28 when it supports the repository's declared architecture. An architecture available only in our older runtime uses that runtime, with a warning in the wizard. Existing deployments keep the runtime they were created with.
Qwen3.8-Flash-Next uses a separate vLLM 0.30 runtime. Inspect
Qwen/Qwen3.8-Flash-Next-FP8 for the official FP8 checkpoint,
or Qwen/Qwen3.8-Flash-Next to keep the original BF16 weights.
On the hardware step, choose GPU only or
GPU + RAM. Selecting RAM offload does not change precision.
GPU + RAM keeps the model's PLE lookup tables in pinned host RAM and fetches the rows needed for each token. The MoE and dense weights still occupy GPU memory. The audited FP8 revision moves about 47.7 GiB to RAM; BF16 moves about 95.4 GiB. These estimates apply to the pinned official revisions, not arbitrary forks or future checkpoint updates.
| Checkpoint | GPU + RAM candidates | Available host RAM required |
|---|---|---|
| Official FP8 | RunPod 2 × H100; Nebius 1 × B200 | 128 GiB minimum |
| Original BF16 | RunPod 2 × H200 or 4 × H100 | 192 GiB minimum |
These are sizing estimates, not throughput benchmarks or capacity
reservations. RunPod placement requests RAM above the required floor.
Nebius's B200 option has 224 GiB RAM and runs in us-central1,
so it is unavailable with EU-only residency. Its compute budget uses
$8.50/hour conservatively, covering the announced October 1 price increase
from $7.15; the model disk is billed separately. Each provider needs its
own linked account.
On September 23, 2026, the official FP8 checkpoint served a successful chat response on two RunPod H100s with 503 GB of host RAM. The runtime confirmed pinned CPU storage for the PLE tables and reported 62.46 GiB per GPU for model loading, before KV cache and serving overhead. This validates that configuration; the other combinations above remain sizing estimates.
The runtime checks available RAM inside the container before downloading weights. Swap does not satisfy this requirement. A small 24 or 48 GB GPU still cannot hold the remaining FP8 model weights with this offload method. Create a new deployment to use the new runtime and memory mode.
Before provisioning, we check the declared quantization method, common
AWQ, GPTQ and FP8 settings, activation dtype, and GPU parallelism against
the model configuration. These checks catch known incompatibilities; they
do not guarantee that every checkpoint or GPU kernel will work. For
Altworld/Hemmingway-1, start with one 80 GB GPU and create a
new deployment to use the compatible runtime.
Pick a shape and region
A shape is the GPU configuration the model runs on, for example one H100 or two H200s. The wizard shows shapes that fit the model, with an hourly and monthly price for each. Regions come from your linked provider. If your data must stay in the EU, pick an EU region: the deployment and its storage stay there.
Choose when it runs
In Configuration, enable Schedule this deployment to choose weekly running windows before deploying. Presets cover office hours, all weekdays and overnight runs. Set your time zone, add windows and review the hours and savings preview. Pre-warm can be automatic or a lead you choose.
The first deployment provisions immediately. Once ready, it stops outside its windows and wakes ahead of later windows. The budget cap and deletion timer still apply; the timer counts serving time across wakes. For a recurring service, consider turning that timer off. See schedules for costs, overnight windows and manual holds.
Review and confirm
Before anything is provisioned you see a summary: the model, the shape, the region, the estimated hourly and monthly cost, and any schedule you configured. The schedule is saved with your confirmation and activated together with the deployment. The prices shown are estimates from the catalog, not a quotation. Your cloud provider bills you directly for all infrastructure usage, which is why the confirm step asks you to acknowledge exactly that. Set a budget cap here too.
What happens during a boot
After you confirm, LLM Hangar provisions the instance in your account, configures the inference server (vLLM or SGLang), downloads the model weights, and brings the endpoint up. You can watch each step live in the dashboard. Large models take time on the first boot: weights for a big model are hundreds of gigabytes, so allow time for the initial download and engine preparation. The model guide and dashboard show estimates or dated measurements where available.
Where the provider supports it, LLM Hangar keeps a prepared stage with the weights already in place after a destroy. Redeploying the same model then reuses the stage and boots much faster. The stage's storage lives in your account and is shown with its own cost; you can delete it at any time.
When it is ready
A ready deployment shows its endpoint URL, for example
https://abc12345.gw.llmhangar.com, with a valid certificate,
plus an API key shown once. The Connect tab has copy-paste snippets for
common clients: see using your
endpoint.
Stop, resume, schedules
Deployments can be stopped and resumed without destroying them, and a wake/sleep schedule can keep one running only during working hours, which cuts the infrastructure bill accordingly. Deleting a deployment tears everything down in your account, and the teardown is verified: we check nothing is left running before reporting it destroyed.