LLM Hangar / Private LLM hosting
Private LLM hosting in your own AWS, Nebius, RunPod or Verda account
LLM Hangar deploys an open model on a GPU instance in your cloud account and chosen region. You control the resources and provider relationship; we handle deployment through the access you grant and record the actions in your audit log.
What "private" means here
With the default node-local gateway, requests travel directly between your client and the endpoint on your instance over HTTPS. A private edge also runs in your account. If you choose the hosted edge, prompts and responses pass through a gateway we operate. See gateway placement and access and the sub-processor list.
The model weights and serving process run in your cloud account. Your own client, logging and external tools also affect where data is sent or retained.
Each deployment supports named keys with individual limits. Keys are shown once when created. Key changes reach an edge gateway within seconds; on the default gateway, they take effect at the next start.
The security page explains cloud access, credential storage and teardown checks.
What you get
- An OpenAI-compatible endpoint with a valid certificate. Existing Python, JavaScript, LangChain, Vercel AI SDK and editor configurations work by changing the base URL (snippets).
- A hard budget cap on every deployment. You can raise it; you cannot disable it. At the cap the deployment is destroyed and you get an email naming it and the amount.
- A self-destruct timer, so a forgotten GPU cannot run all weekend. Trial deployments always carry one.
- Verified teardown: after every destroy we check with your provider that the resources are actually gone, and keep checking until that is confirmed.
- An audit log of every action taken in your account, timestamped, including the plan recorded before anything is created. It exports as JSON.
- An EU-only checkbox that pins every resource of a deployment to EU member-state regions and keeps it there.
What it costs
Two bills. LLM Hangar is a flat $39 a month on the Lab plan (€34 in EUR), with a 7-day free trial that requires no credit card. Your cloud provider bills you directly for the GPU hours, at your own provider's rates, with your own credits and reserved capacity if you have them. We never see that invoice.
Two reference points for the GPU side, both cited rather than estimated:
- On AWS, a g6e.xlarge (one NVIDIA L40S, 48 GB) was listed by Spare Cores on 2026-08-22 at $1.974 per hour on-demand and $0.604 spot in Stockholm, and $2.327 / $1.747 in Frankfurt. That shape runs models up to roughly 27B parameters at 4-bit.
- On RunPod, the 2x H200 shape that runs DeepSeek V4 Flash cost $7 to $14 per hour at our own boots in August 2026, the rate fixed at pod creation. The full numbers are in DeepSeek V4 Flash: GPU requirements, boot time and cost.
The wizard shows the exact hourly and monthly estimate for the shape and region you pick before you confirm anything, and a cost calculator lets you compare that against a per-token API at your own volume.
Who it is for
- Teams that need control over where inference runs. Client data, source code, medical or legal records: the prompt stays on a machine in an account you control.
- EU teams who need every resource in EU regions with a clear account of gateway placement and data handling. See EU-hosted LLM API providers compared.
- Solo builders and small teams who want a private endpoint for an evening's prototype and do not want to write Terraform, build an AMI or configure an inference engine to get it.
Limits to consider
- The Lab plan runs one GPU shape per deployment, as listed on the plan card. The largest catalog models need multi-GPU shapes; the per-model guides state which.
- LLM Hangar does not currently hold a SOC 2 or ISO 27001 certification. The controls are described so you can evaluate them directly; the audit log and your own console let you verify them.
- Budget caps and timers are enforced against estimates from catalog prices. Your provider's meter, not our estimate, is the invoice. Caps bound the window in which spend can accumulate; they are not a guarantee of the exact amount.
Frequently asked questions
Is this the Private LLM app?
No. Private LLM is an iOS and Mac app that runs small models on your device. LLM Hangar is hosting infrastructure: it deploys open models onto GPU instances inside your own AWS, Nebius or RunPod account and gives your applications a private, key-authenticated, OpenAI-compatible endpoint.
Do my prompts ever reach LLM Hangar?
With the default gateway, requests go directly to your instance. The optional hosted edge proxies requests through infrastructure we operate. Gateway placement and any audit-capture settings determine the request path and retention; review these before sending sensitive data.
Can I keep the infrastructure if I cancel?
Yes. The instance, disk, network and endpoint are resources in your own account, tagged and owned by you. They stay after you cancel; you can manage them in your own console. Access for LLM Hangar is granted through a role or key you can revoke at any time.
Which providers and regions?
AWS, Nebius, RunPod and Verda. Regions come from your linked provider. The EU-only option pins every resource of a deployment to EU member-state regions, for example eu-central-1, eu-west-1, eu-west-3, eu-north-1, eu-south-1 and eu-south-2 on AWS.