Deploy your own private LLM, on infrastructure you control.
Run an open model in your AWS, Nebius, RunPod or Verda account. LLM Hangar sets up the GPU server and gives your applications an API endpoint protected by a key.
No credit card required · one deployment at a time
Pick an LLM and deploy it
to your GPU provider
to your GPU provider
Set running hours,
budget caps and timers
budget caps and timers
Delete a deployment
and check its cleanup
and check its cleanup
Choose where
your model runs EU
your model runs EU
No command line or Linux experience needed. No Terraform to write, machine images to build or inference engines to configure.
Private LLM hourly costs
Use these prices to estimate what a model costs to host per month, and what your own machine can run.
Loading…
Loading current prices…
Manage resources in your own account
The instance, storage and network are created in your account. You can inspect and manage them in your provider console, apply your own policies and retain them if you cancel your subscription.
No lock-in: resources stay after you cancel
Your provider rates, your reserved capacity, your credits
Access granted by a role you can revoke in one command
WHEN
WHAT
WHERE
08:03:08
Plan recorded: 1 instance, 1 security group
us-east-1
08:04:20
GPU instance created, billing starts
g6e.xlarge
08:19:12
Endpoint live and answering
203.0.113.10
21:44:03
Destroyed and swept, zero resources remaining
us-east-1
See every step of your deployment
You see what ran, where, when and how. Every action we take in your account is written to a timestamped log, including the plan of what will be created before it exists.
The log exports as JSON, is kept after teardown, and sits alongside spend against cap and request counts read from your own gateway.
Choose how requests reach your model
A dedicated model deployment
Model weights are downloaded to the instance in your cloud account. Your deployment has its own serving process and endpoint.
Direct requests by default
With the default gateway, requests go directly to your instance. You can also choose a private gateway in your account or a hosted gateway that proxies requests through our infrastructure.
EU-only infrastructure
EU
One checkbox pins every resource to EU member-state regions and keeps it there.
eu-central-1
eu-west-1
eu-west-3
eu-north-1
eu-south-1
eu-south-2
Access you control
Create named endpoint keys with their own limits. Keys are shown once; changes take effect within seconds on an edge gateway or at the next start on the default gateway. You control the cloud credentials that let us manage your deployment.
AVAILABLE CONTROLS
EU region selection
Direct request routing by default
Encryption in transit
Full audit trail
Verified deletion
Pricing
One flat subscription for the platform. GPU hours are billed to you by your cloud provider.
Launch discount
50% off
Lab
Evals, prototypes and side projects.
$39
/month*
$39
Hardware1 GPU
Deployments at once2
Modelswhole catalog
ProvidersAWS, Nebius, RunPod and Verda
History after teardown30 days
Supportemail
7-day free trial, no credit card required.
* The subscription covers the deployment platform. GPU and hosting charges are billed by your infrastructure provider, like AWS or RunPod, directly to you.
Included on every plan, including the trial
Hard budget caps
Required on every deployment. You can raise the cap; you cannot disable it.
Weekly wake/sleep schedules
Choose running windows while deploying, with time zones, overnight runs, pre-warm and a savings preview. How scheduling works.
Self-destruct timers
Set when a deployment should be deleted, even if you are away.
Verified teardown
A sweep after every destroy, with a record of what was removed.
Full audit trail
Every action we took in your account, timestamped, including the pre-flight plan.
EU-only residencyEU
One checkbox pins a deployment to EU-member-state regions.
Delete
Delete any deployment, or your whole organization, at any time on any plan.
Frequently asked questions
How does the 7-day free trial work?
The trial runs for 7 days and no credit card is required to start it. You get the Lab feature set with one deployment at a time. Subscribe whenever you are ready; if the trial ends first, new deployments are blocked until you do. Nothing is destroyed: anything still running keeps running and keeps billing to your cloud account.
What does a stopped deployment cost?
GPU billing ends once the provider confirms the compute is off. Retained storage and an optional private gateway can still cost money. The platform subscription stays the same. Your configuration and keys are kept for the next start. See stop and start.
Do you resell GPU capacity?
No. Everything runs in your account at your provider's rates. We never see your infrastructure invoice.
What access do you need to my cloud?
On AWS, a role we can borrow for an hour at a time, scoped to resources we tagged ourselves. Delete the CloudFormation stack to remove the role; any session already issued can remain valid until it expires.
What if a deploy fails halfway?
We explain the cause in plain language, remove whatever was created, and show the sweep result. The most common first failure is an AWS GPU quota of zero, which comes with a direct link to the form that fixes it.
From the blog
All posts →Spot vs on-demand GPUs for LLM inference | LLM Hangar
Decide whether spot GPUs suit your LLM workload. Calculate useful-work cost, account for interruptions and test recovery before relying on the discount.
Qwen 4 announced: what we know so far | LLM Hangar
Alibaba has announced Qwen 4, including Max, Plus, Flash and 27B. Here is what was confirmed, and what is still unknown about the release.
A private AI for your team: size it and price it
Pick a model, add your people and find a starting setup. Estimate cost per person and check what changes when everyone prompts at once.
Nemotron's IOI 2026 gold: 760 GPUs, 1,000 tries
Nvidia's Nemotron beat the top human at IOI 2026. The paper's own numbers show the model answering once scored 304, and what the rest of the score cost in GPUs.
Is Ox Alpha open weights? Yes: it is GLM-5.3-Flash
Ox Alpha was revealed as Z.ai's GLM-5.3-Flash. Find the official MIT-licensed weights, check a conversion's provenance and understand the hosting options.
GLM-5.3: hardware requirements and license
GLM-5.3 weights are available. Check the FP8 and BF16 hardware requirements, current license and serving setup before renting a GPU node.
DeepSeek V4 Flash: GPU requirements, boot time and cost
Our DeepSeek V4 Flash deployment on 2x H200: measured cold boot, historical GPU costs, memory requirements and how to assess an API alternative.
Set up your first private model.
No credit card required. Connect a cloud account and pick a model.