Documentation

LLM Hangar deploys open large language models onto GPU infrastructure in your own cloud account and gives you a private, key-authenticated, OpenAI-compatible endpoint. You can manage the resources in your provider console. Your provider bills you directly under your own agreement.

Quickstart

  1. Create an account and start the 7-day free trial. No credit card is needed.
  2. Connect a cloud account: AWS, Nebius, RunPod, or Verda.
  3. Deploy a model from the catalog. Pick a GPU shape and region, review the cost estimate, confirm. For typed decisions instead of text, see decision models.
  4. When the deployment is ready, open its Connect tab for the endpoint URL and API key.
  5. Point any OpenAI-compatible client at it.

Guides

Connect your cloud account

AWS via a CloudFormation cross-account role, Nebius via a project-scoped service account, RunPod via an API key, Verda via client credentials. What each needs and what we verify.

Deploy a model

The deployment flow from catalog to running endpoint: shapes, regions, cost estimates, and what happens during a boot.

Use your endpoint

Working snippets for Python, curl, JavaScript, LangChain, the Vercel AI SDK, editors, and automation tools.

Decision models

Deploy Laya for typed choice, score and yes or no answers with probabilities, and call it with the TypeSafe clients, curl or any HTTP tool.

Budget caps and auto-destroy

How budget estimates, deletion timers, schedules and teardown checks help you manage costs.

Stop and start

States including stop_failed, the deploy-time staging choice and its cost, wake estimates, what a stopped endpoint answers, and the API.

Schedules

Configure weekly windows during deployment or later, with overnight runs, time zones, savings previews and pre-warm, holds and wakes, webhook deliveries and signature verification.

Spot instances

Eviction per provider, automatic recovery from staged weights, the fallback and wait policies, and reacquiring spot.

Keys, sign-in and audit capture

Named keys with limits, JWT validation against your identity provider, and audit capture to a bucket you own with chains you can verify.

Connect Verda

Connect your Verda project and manage instances, staged weights and startup scripts.

Guide: deploy an LLM on AWS without the CLI

Which EC2 instance fits which model, prices by region with their sources, the quota and capacity blockers, and the four steps.

Security: access, keys and data

Where prompts go, how access to your cloud account is scoped, credential storage, the audit log, verified teardown.