LLM Hangar / Models
Models
The catalog pairs open-weight models with suitable GPU configurations. Measured entries identify configurations we have booted; other entries are estimates. You can review the estimated cost before creating resources. Prompts and responses go straight from your client to the endpoint on your instance. If the model you want is not listed, you can also deploy a model from a Hugging Face repository by pasting the repo id and picking a shape yourself.
Catalog and current prices
The table is read live from the same public price feed the homepage uses. Every eight hours we check which shapes are rentable at each provider and at what rate. A plain figure comes from a configuration we checked was deployable; a figure marked with ~ is an estimate from the provider's current GPU rates. Your provider bills you for the GPU; the prices here are not a quotation.
Guides
The guides explain memory requirements, deployment choices and costs. Where we have measured a deployment, they give the configuration and date so you can judge how the result applies to your workload.
- DeepSeek V4 Flash: GPU requirements, boot time and cost: 284B mixture of experts, 2x H200, measured cold boot.
- GLM-5.3: hardware requirements and license: weights available; current FP8 hardware guidance and the flagship's license.
- Gemma 4 31B VRAM requirements: why the weights fitting is not enough: measured VRAM requirements, the real context ceiling on 80 GB cards, and the w4a16 build.
Kimi K3 appears in the price index, but its large checkpoint requires a carefully matched multi-GPU configuration. Treat memory estimates as an initial check and verify the current serving recipe before renting a node.
How a deployment works
- Connect a cloud account: AWS, Nebius, RunPod or Verda.
- Pick a model and a shape from this catalog; choose a region, or tick EU-only.
- Set a budget cap and confirm the estimate.
- Point any OpenAI-compatible client at the endpoint, as in using your endpoint.