Blog / Model guides
Is Ox Alpha open weights? Yes: it is GLM-5.3-Flash
Yes. Ox Alpha was the preview name for Z.ai's GLM-5.3-Flash. You can download the official weights from zai-org/GLM-5.3-Flash, which lists an MIT license. You do not need to search for a repository called Ox Alpha.
OpenRouter's preview listing now identifies Z.ai as the developer and links to GLM-5.3-Flash. That resolves the identity question that motivated the original version of this article.
What changed after the preview?
Ox Alpha appeared as an anonymous hosted model in August 2026. Users could try its coding and reasoning capabilities, but a working API did not establish who made it or whether weights would follow.
The named release makes two practical things possible: downloading a checkpoint from the developer and reviewing the license attached to it. Those are stronger foundations for a deployment decision than similarities between a preview's answers and another model's behavior.
| Question | Where to check |
|---|---|
| Who made the preview? | OpenRouter's identity notice on the Ox Alpha listing |
| Where are the official weights? | Z.ai's GLM-5.3-Flash repository on Hugging Face |
| Which license applies? | The license in the exact repository and revision you download |
| What does self-hosting require? | The current engine recipe and the checkpoint's actual size |
Check a converted checkpoint before using it
A repository name is not proof of provenance. For a GGUF or another community conversion, look for the original model ID, source revision, conversion method and any quality checks. Verify that weight files are actually present; a model card or configuration file alone is not a downloadable model.
This matters beyond the Ox Alpha release. A conversion can use an older checkpoint, change the chat template or require a different engine. Save its revision in your deployment record so a later update does not silently change the model you tested.
Open weights and open source answer different questions
Open weights means you can obtain the parameters needed to run the released model. It does not by itself give you the training data and complete process needed to reproduce it. Describe GLM-5.3-Flash as an open-weight model with an MIT license when those are the properties you mean.
Also keep Flash separate from GLM-5.3: hardware requirements and license. The flagship has a different checkpoint and its own license. Approval to use one release should not be copied to the other without checking.
Can you run it yourself?
Our catalog's native FP8 checkpoint is about 328 GB, and we measured a deployment on four H200 GPUs. It took about 34 minutes to reach a ready endpoint on 27 August 2026.
For an occasional experiment, compare the cost of that node with the hosted API before renting it. For a private service, test your exact client and workload: streaming, image inputs, tool calls and long conversations may place different demands on the engine.
Review the request path as well as the weights
The old preview listing says the provider retained prompts and completions without using them for training. That is a description of the preview's policy. Check the current provider's terms when choosing a hosted endpoint.
Self-hosting lets you choose where inference runs, but your editor, application logs, monitoring and gateway still determine where data travels. With LLM Hangar's default node-local gateway, requests go directly to your instance; the optional hosted edge introduces a managed proxy. The access guide explains those choices.
Frequently asked questions
Who made Ox Alpha?
Z.ai. OpenRouter's preview listing now identifies the model as GLM-5.3-Flash and links to its production entry.
Where should I download the weights?
Start with zai-org/GLM-5.3-Flash on Hugging Face. For a community conversion, check its source model, revision and conversion method before using it.
Is downloading the weights enough to make my application private?
No. You also need to review your client, gateway, logs and external tools. They can send or retain data even when inference runs on your own instance.