LLM Hangar / Team planner
A private AI.
A budget for your whole team.
Pick a model. Add your people. Find a starting setup and its monthly cost, with a clear check on simultaneous use.
Can 25 people prompt it at once?
Yes, a serving engine can handle requests from multiple people, batching work and queuing requests when necessary. The useful question is whether your chosen model and GPU can answer all 25 quickly enough. This planner checks the memory side and gives you a speed target to verify.
People ≠ simultaneous answers
A team of 25 may have only a few people asking at once. A workshop where everyone presses Enter together is a different workload. Use “At the same time” to plan for that peak.
The model shares; chat memory grows
The model is loaded once. Each active conversation needs its own working memory. Long documents and conversation history increase that memory, so the same GPU can fit fewer active chats.
Fitting is only the first check
A GPU may hold every conversation while answering too slowly. Larger queues let requests wait; they do not create more computing power. Test response time at your busiest moment.
Turn this estimate into a test
For you or the person setting this up: test the selected model, build and GPU together. Start small, then run the peak below. Record first-response time, answer speed and errors. A successful single chat does not prove multi-user capacity.
How the memory estimate works
We start with the selected model build’s weight files, reserve 15% of total GPU memory for the engine and overhead, then estimate the extra memory for each active conversation. This is a planning estimate, not a measured maximum. Engine settings, memory fragmentation, cached prefixes and request length can change the result.
Estimated active chats = (GPU memory × (1 − reserve) − model weights) ÷ memory per chat, rounded down. We cap this at the catalog’s explicit request limit, or use 256 as a planning ceiling when no limit is recorded. A request must also fit that setup’s configured context length. GPU GB is treated conservatively as decimal GB.
The smaller Qwen models use their model-specific attention configuration. Gemma uses a conservative recorded cache-allocation estimate. The large hybrid models need a measured cache budget before we can estimate simultaneous chat memory.
Short chats allow 1,536 input and 512 output tokens; documents allow 7,168 + 1,024; coding allows 14,336 + 2,048. Input includes the whole conversation and instructions. Output includes any reasoning tokens. These are example workloads, not limits on the kinds of work a model can do.
See vLLM’s memory and batching guidance and hybrid cache allocation. No controlled multi-user throughput benchmark is available in this planner’s catalog snapshot; all speed values shown are targets.
What is included in the monthly cost?
GPU rate × (serving hours + startup hours), plus retained storage, the platform fee and any extra costs you enter. Weekdays assume 22 working days. Always-on uses 730 hours, including startup. Turning the GPU off cuts compute charges; retained storage and the subscription can keep billing.
The default $39/month uses LLM Hangar Lab, which covers one GPU. Multi-GPU setups need a platform quote; their displayed amount is a subtotal until you enter it. Startup times and storage are estimates, not guarantees for every shape or region. Storage budgets cover retained model files; additional disks, network traffic, backups and operator time must be added under other costs. Tax is excluded. All amounts are USD.
Prices come from the public GPU price index or dated catalog quotes. A list estimate is not a reservation or confirmation of availability. A saved link keeps the selected hourly quote; “Choose for me” returns to current estimates. Model and hardware configurations were reviewed on 10 September 2026.
Already know your hardware? Try the hosting cost calculator. Comparing token bills? Open the API break-even calculator.
Published · Updated