Blog / Model guides
GLM-5.3: hardware requirements and license
GLM-5.3's weights are available on Z.ai's official repository. The main release uses FP8, while BF16 has a separate repository. This replaces the release-date estimate in the earlier version of this guide.
For a private deployment, start by checking the exact checkpoint and serving recipe. We have not measured a GLM-5.3 deployment for this article, so the configurations below are documented starting points, not LLM Hangar benchmark results.
Which hardware should you plan for?
The vLLM recipe checked on 6 September 2026 uses an eight-H200 node for the standard FP8 setup. It gives a separate eight-B200 configuration for the full one-million-token context window and points BF16 users to a larger, multi-node setup.
| Checkpoint or workload | Planning guidance |
|---|---|
| Official FP8 release | Start with the recipe's 8x H200 configuration; validate your own context and concurrency |
| Full 1M context | The recipe uses 8x B200 with FP8 KV cache; loading the weights alone does not establish context capacity |
| Official BF16 release | Use the separate GLM-5.3-BF16 repository and plan the topology before renting hardware |
| Community quantization | Check provenance, required GPU architecture and quality on your tasks |
An eight-H200 node has 1,128 GB of advertised aggregate GPU memory. How much remains usable for requests depends on the engine, weight distribution, cache format and runtime allocations. Treat aggregate memory as a first screening step. It does not prove that a particular launch configuration will work.
For disk sizing, inspect the files in the chosen repository and allow room for the container image, cache and temporary files. Avoid estimating a mixed-precision checkpoint solely by multiplying a rounded parameter count by one byte.
The flagship has its own license
GLM-5.3 uses the GLM-5.3 License, rather than the MIT license attached to GLM-5.3-Flash. It broadly permits use and modification, subject to its conditions.
One condition concerns organizations that operate a model-as-a-service business, including through affiliates, and whose combined revenue exceeds US$10 billion over any consecutive 12 months. They must pass Z.ai's security review before commercial use. The license defines that business category and excludes certain embedded end-user features and simple request relaying. Retain the notices and review the actual text for your organization and intended service.
Record the model revision and license alongside your deployment configuration. That gives the next reviewer a clear answer to which release was approved, instead of leaving them to infer it from a changing repository name.
Make the first evaluation small enough to learn from
Start with a short, bounded trial of the workload you actually want to serve. A model's coding benchmark can justify investigating it, but it cannot tell you how your editor, tools or application will behave.
- Choose a few representative tasks with answers you can check: a code change, a document question and a tool call, for example.
- Set a realistic context limit and output budget. Record the reasoning setting explicitly so later runs are comparable.
- Check time to first token, total response time, correctness and tool-call formatting.
- Repeat under expected concurrency. Inspect memory use and queues before raising limits.
- Stop the node and confirm teardown after collecting the results.
The model card describes configurable reasoning effort. Try the setting that meets your latency target, then check whether reducing effort changes the success rate on your tasks. A faster answer that needs a second attempt may consume more time overall.
Work out the cost before starting
Use the price of the whole node, not one GPU. As an illustration, a node quoted at $40 an hour costs $160 for four billed hours or $28,800 for a 720-hour month, before storage and other fees. Those are example numbers, not a provider quotation.
Allow time for downloads and engine startup inside the evaluation budget. If you will repeat the evaluation, compare the storage cost of retaining weights with the time saved on the next boot. The hosting calculator can help with that estimate; check its assumptions against the exact checkpoint and live provider offer.
How GLM-5.3-Flash differs
GLM-5.3-Flash is a separate model with different hardware requirements and an MIT license. Its hardware requirements, boot time and cost do not transfer to the larger flagship.
Frequently asked questions
Are the GLM-5.3 weights available?
Yes. The official Z.ai repository carries the FP8 release as of this 6 September 2026 review. BF16 is published separately.
Is GLM-5.3 MIT-licensed?
No. The flagship has the GLM-5.3 License, including a security-review condition for certain large model-as-a-service businesses. GLM-5.3-Flash has a separate MIT license.
Have you measured its boot time?
No GLM-5.3 flagship boot measurement is published in this guide. Its hardware guidance comes from the current serving recipe, and should be validated for your workload.