Four days ago I set a test for Z.ai. If the GLM-5.3 weights shipped on time and stayed free for commercial use, the two-week delay was a safety pass. If the date slipped or the terms tightened, it was something else.

On August 28, fourteen days to the day, the file appeared. The terms tightened. Both halves came true at once, and the way they did is more interesting than either.

What Shipped

  • August 26. GLM-5.3 Flash: 320 billion parameters with 18 billion active, a new base model, 1.3 million tokens of context. Licence: MIT. Price: $0.15 per million input tokens, $0.50 output, half that during launch.
  • August 28. The GLM-5.3 weights, two repositories on Hugging Face, an FP8 build and a BF16 build one Hacker News commenter put at 770GB. The thread ran to 795 points.

The benchmark table on the release is the benchmark table from August 14. CyberGym 84.5, ExploitBench 54.4, DeepSWE 66.9. Not a decimal moved.

The Flash Model Was Already Out

The week before the Flash announcement, a free model called ox-alpha appeared on OpenRouter and OpenCode. People on r/LocalLLaMA ran tokeniser fingerprints against it and called it as GLM-family before Z.ai said a word.

On August 26 Z.ai confirmed it. Ox-alpha was GLM-5.3 Flash, tested anonymously “to gather user feedback.” It “quickly became the most popular model of the week,” the post says, “with all of this traffic served on Chinese AI chips,” at “per-token cost comparable to mainstream NVIDIA GPUs.” Take that as a vendor claim. But notice the shape: the cheap model went out under an alias, was load-tested by the public for free, and only then got a name and a licence. The expensive model went out with a name, a two-week hold, and a licence with a clause in it.

The Clause

The GLM-5.3 licence reads as MIT until you reach the definitions. “Model as a Service” means giving a third party access to inference or fine-tuning with meaningful control over inputs, parameters or training data. Embedding the model in a product feature does not count. Relaying requests to someone else’s hosted copy does not count.

Then the condition. If a licensee and its affiliates have combined revenue above ten billion US dollars over any consecutive twelve months, they must complete Z.ai’s security review before commercial use.

The scope and method of the security review shall be reasonably determined by Z.AI.

— GLM-5.3 licence

Work through who that catches. Amazon, Google, Microsoft, Oracle, Nvidia. Alibaba, Tencent, ByteDance, Huawei. Every company that can serve a model this size at real scale is over the line. Fireworks, Together and the quantisers on r/LocalLLaMA are nowhere near it.

Meta wrote the template. The Llama licence has required a separate agreement from anyone with more than 700 million monthly users since 2023, and nobody calls Llama closed for it. Z.ai swaps the commercial licence for a security review, and that swap is the whole story. A Chinese lab now holds a standing right to inspect the world’s hyperscalers before they can sell its model. The criteria are not published. The lab decides what “reasonable” means.

Open weights now has a revenue test

The file is open. Anyone can download and run it. What changed is that “open” and “sellable by anyone” have been separated by a number, and the number sits just above every company that could turn this model into a mass-market API overnight.

What the Fourteen Days Were For

The August 14 post said the weights would follow “once safety evaluation and hardening are complete.” What was hardened is still not explained. The cyber scores at release match the cyber scores at launch. Either the hardening did not touch measured capability, so it was refusal training or nothing, or the numbers were not rerun. Neither is what “hardening” was meant to suggest.

Meanwhile the Flash benchmark table has no cyber rows at all. The model that shipped MIT is the one not marketed on hacking. The model marketed on hacking has the review clause. Tidy policy or smaller table, I cannot tell from outside.

What I can say is that the hold and the clause do the same job by different means. The hold controlled who could run the model for two weeks. The clause controls who can sell it, forever, at the lab’s discretion. The safety mechanism did not end on August 28. It changed form, from a date into a contract.

The Third Option

In July I argued Washington could not ban the weights because the only window closes on release. Last week I argued the publisher is the whole mechanism. This release shows an option I did not list: publish, and keep a lever anyway. Not over the file, which is gone, but over the money. Anyone can run GLM-5.3. The people who could make it the default for a hundred million users cannot do so without asking first.

One live consequence. Nvidia is reported to be buying Hugging Face. If that closes, the site hosting the weights becomes an affiliate of a company over the line. A question, not a ruling, but the sort of question the clause was written to raise.

What To Watch

  • Whether any hyperscaler serves GLM-5.3 at all. If Bedrock and Vertex list Flash and skip the flagship, the clause worked exactly as designed.
  • Whether the review criteria are ever published. Until they are, “reasonably determined” is the entire policy. The August 14 page, as of this writing, still says the weights are coming in two weeks.

Last week’s conclusion was that the entire safety system for open models is one company’s product decision. That has not changed. What changed is that the decision now has a legal instrument attached, pointed at the companies most able to put a frontier model in front of everyone. The summer’s argument was about the file. The file shipped. The argument moved into the licence, and almost nobody reads those.