GLM-5.3-Flash runs locally — and it's MIT licensed
In August 2026, a model called "Ox Alpha" appeared anonymously on OpenRouter and OpenCode. In five days it served roughly 50 trillion tokens of traffic. Nobody knew who built it — until tokenizer forensics and error-code fingerprints pointed at Zhipu, and Z.ai confirmed it to Bloomberg.
Ox Alpha was GLM-5.3-Flash: Zhipu's fast-tier flagship, officially launched August 26, 2026. And the weights are MIT licensed.
The specs
320B total parameters, 18B active — a sparse mixture-of-experts. 1M-token context, 128K max output, natively multimodal (text, image, video), pretrained on 30T tokens. On agentic coding benchmarks it trades blows with Claude Opus 4.8: it leads on GDPVal-AA v2 (1773 vs 1582) and DeepSWE v1.1 (63.4 vs 58.0). The API costs $0.15/$0.50 per million tokens — roughly a tenth of prior frontier-tier pricing.
But the API isn't the interesting part. The weights are public, ungated, on Hugging Face
(zai-org/GLM-5.3-Flash), under the MIT license. Commercial use allowed. No gatekeeper.
The honest sizing math
320B parameters at BF16 is about 642GB — not fitting on anyone's desk. But MoE is the local-inference cheat code: only 18B parameters are active per forward pass. You load the library; you only read the shelf you need.
Quantized, the math works out:
- 4-bit: ~160–180GB — server territory
- 3-bit: ~128GB — tight on a 128GB box
- 2-bit: ~101GB — comfortable
- 1-bit (Unsloth dynamic): ~93GB — genuinely fits a 128GB desktop, retaining about 71% top-1 accuracy
llama.cpp merged full support (hybrid text+vision) on September 30, 2026. Real-world: a 2-bit 101GB quant decodes at 73.5 tok/s on 2× RTX PRO 6000 Blackwell. A frontier-adjacent 320B MoE, running on a desktop, at interactive speed.
Why this matters for security work
Client engagements produce the most sensitive data in the industry: proprietary source code, vulnerability findings, packet captures, infrastructure maps. Every API call ships that data to a third party. Contracts get signed, DPAs get negotiated, and everyone quietly hopes the vendor's retention policy means what it says.
MIT frontier weights on local hardware end that conversation. The model runs on a box you own. Nothing leaves the building. No DPA, no retention policy, no trust required — the data physically cannot go anywhere.
That's not a minor operational detail. For pentest firms, dev shops handling client code, and anyone running evals against proprietary systems, private inference is a genuine differentiator — and as of this year, it no longer means accepting a weak model to get it.
If your security workloads are stuck choosing between capable models and data privacy, that tradeoff just expired. Book a scoping call and we'll map what a private pipeline looks like for your stack.
Need a pentest, an AI security assessment, or a custom security build?
Human-led testing, production AI builds, and the full loop in between. Book a free 30-minute scoping call.
Book a scoping call