Mystery model 'Ox Alpha' confirmed as GLM-5.3-Flash; independent quantization and serving benchmarks emerge
2026-08-27
Per Latent Space's AINews roundup, the previously unidentified model 'Ox Alpha' was confirmed to be Z.ai/Zhipu's GLM-5.3-Flash (320B total params, 18B active, 1M context, hybrid attention). Unsloth reported the model runs as 3-bit GGUF on 128GB RAM, with 4-bit retaining 93% accuracy and fitting a 256GB Mac or two DGX Sparks. Baseten reported 122+ TPS serving throughput on day 0; Databricks cited 270 tok/s and 10% higher quality than GLM-5.2 at 1/10 the cost on a benchmark it calls 'OfficeQA Pro v2.' Together AI said it nearly matches a model it calls 'Luna' on a benchmark called 'DeepSWE' at less than half the compute budget.
Significance 4: Independent quantization work and serving benchmarks from multiple named infra providers (Unsloth, Baseten, Databricks, Together AI) corroborating an open-weights model's price-performance claims moves this beyond lab-reported status, meeting the bar for moving a tracked axis (cost, coding/agentic) on the open-weights frontier.
Implications · machine-drafted, not owner judgment
Independent (not just lab-reported) confirmation of strong intelligence-per-dollar for an MIT-licensed-class open model, plus quantization recipes making it deployable on prosumer hardware (256GB Mac, two DGX Sparks), extends the open-weights frontier and directly bears on the bulk-cheap model verdict, where cost/speed are the deciding axes. The 'OfficeQA Pro v2' and 'DeepSWE' benchmarks are not on the tracked known-methods list, so their comparisons should be treated as unverified until reproduced on a tracked method.
- A tracked benchmark (SWE-bench Verified, Artificial Analysis Intelligence Index) picking up GLM-5.3-Flash
- Whether GLM-5.3-Flash's price-performance holds up against DeepSeek V4 on the bulk-cheap verdict's evidence axes
Sources
- curator [AINews] OpenAI to reach AGI bar by end-2026 retrieved 2026-08-28