Z.ai open-weights the full GLM-5.3 model (744B/40B active), positioned for agentic coding and cyber defense
2026-08-28
Per Latent Space's AINews roundup of AI Twitter activity, Z.ai (@Zai_org) open-weighted the full GLM-5.3 model — distinct from the smaller GLM-5.3-Flash launched days earlier — positioning it for agentic coding and cyber defense. vLLM (@vllm_project) confirmed day-0 serving support, citing 744B total parameters, 40B active, a 1M-token context window, and 128K max output, reusing the GLM-5.2 serving path. Practical local-hardware requirements were summarized as ranging from 10-12x H100 GPUs at FP8 down to aggressive low-bit Mac Studio configurations; Unsloth (@UnslothAI) said a 239GB 2-bit quantized variant retains about 81% accuracy after shrinking from the 1.51TB full-precision model.
Significance 3: A second major open-weights release from Z.ai within days extends the open frontier again, but the only source is a single curator's relay of social-media posts with no primary Z.ai announcement link, so significance is held down pending corroboration.
Implications · machine-drafted, not owner judgment
A second Z.ai open-weights release in the same week, this time explicitly pitched at agentic coding and cyber defense, signals Z.ai is trying to establish a full product line (flash + full) at the open frontier rather than a single opportunistic drop. If the cyber-defense framing holds up, it also puts an open-weights model explicitly in a dual-use security posture, which is worth tracking alongside the separate report this run of AI agents accelerating exploit discovery.
- A primary Z.ai blog post or model card confirming these specs
- Independent SWE-bench Verified or Terminal-Bench numbers for full GLM-5.3
- Any documented use of GLM-5.3 in an actual cyber-defense or offensive-security context
Sources
- curator [AINews] OpenAI shuts off Cursor retrieved 2026-08-29