Zhipu’s release note for GLM-5.3 comprises a sentence that didn’t make it into a lot of the protection. Describing its personal cybersecurity outcomes, the Beijing firm writes that functionality “is rising quickest precisely the place we’re furthest behind.”
Zhipu, which additionally trades as Z.ai, is one in all a handful of Chinese language labs releasing fashions that compete with the American frontier. On August 14, it launched GLM-5.3, a coding-focused mannequin, and printed a technical launch notice setting out how the mannequin performs towards its rivals. That notice is the supply for every thing reported right here.
The declare that travelled was about safety. Alongside the coding outcomes, Zhipu stated GLM-5.3 had grow to be unexpectedly good at discovering software program vulnerabilities, scoring 84.5% on a benchmark referred to as CyberGym towards 83.8% for Anthropic’s Mythos 5 and 83.6% for OpenAI’s GPT-5.6 Sol. Headlines adopted reporting {that a} Chinese language mannequin now out-finds the American ones at bug looking.
The explanation that lands more durable than a standard benchmark result’s what vulnerability discovery has grow to be. A mannequin that may learn a codebase and find exploitable flaws is helpful to a defender auditing their very own software program and helpful to anybody doing the identical to someone else’s. Anthropic’s equal work sits behind restricted entry for that purpose, whereas Zhipu intends to publish GLM-5.3’s weights for anybody to obtain.
Zhipu’s personal launch is extra measured than the protection it produced. The CyberGym quantity is actual, and it’s within the paper. It is usually the narrowest of the three cybersecurity outcomes the corporate printed, and Zhipu is upfront that the opposite two go the opposite method.
Three benchmarks, three completely different footage
CyberGym begins from supply code the mannequin can learn and exams whether or not it might probably discover a vulnerability and make sure the flaw is real. That’s the consequence that travelled, and the margin is seven tenths of a share level.
ExploitBench asks one thing more durable, requiring the mannequin to purpose about an actual vulnerability and the way it will be exploited. GLM-5.3 scores 54.4%, greater than double its predecessor’s 24.4%. Mythos 5 scores 78.0% and GPT-5.6 Sol 76.5%.
ExploitGym counts what number of exploitation duties a mannequin finishes inside a hard and fast time funds. GLM-5.3 completes 105 duties in two hours and 130 in six. Mythos 5 completes 181 and 247.
These two outcomes have been reported thinly, and they’re those that describe the hole. Discovering a flaw and constructing a working exploit from it are completely different jobs. Zhipu’s studying is that the additional alongside that chain a take a look at sits, the additional behind its mannequin is, and the corporate says so within the launch somewhat than leaving it to be found.

Which Anthropic mannequin, and why it retains altering
A part of the confusion within the protection comes from Zhipu evaluating three completely different Anthropic fashions in three completely different locations. The principle benchmark desk units GLM-5.3 towards Opus 4.8. The efficiency charts use Fable 5. The cybersecurity part makes use of Mythos 5. Anybody studying rapidly comes away with a single comparability that doesn’t exist.
On coding, the image is combined somewhat than dominant. GLM-5.3 leads Opus 4.8 on some exams and trails it on others, and Zhipu states plainly that its mannequin stays behind Claude Fable 5 on the corporate’s personal inside coding benchmark.
How the exams had been run
The methodology footnotes comprise one thing the summaries skipped. Zhipu evaluated GLM-5.3 on CyberGym, ExploitGym, ExploitBench, Terminal Bench and a number of other different duties inside Claude Code 2.1.207, Anthropic’s coding agent.
That isn’t improper. Utilizing a standard harness throughout fashions is how a comparability stays honest, and Zhipu paperwork the settings it used. It’s price noticing anyway. A Chinese language open-weights mannequin’s frontier claims are being measured via American agent software program, which says one thing about the place the tooling layer sits on this competitors that the mannequin scores don’t.
Two additional particulars deserve consideration earlier than the CyberGym result’s handled as settled. The rating is a single run, reported as cross@1 throughout 1,507 duties, with no variance figures given. A niche of seven tenths of some extent between two single runs isn’t a niche anybody ought to lean on. And the ExploitGym time budgets had been normalised utilizing throughput charges from Synthetic Evaluation, with rescaling elements listed for GLM-5.3, Kimi K3 and Qwen3.8 Max, however not for Mythos 5.
The vulnerability depend and the quantity that’s lacking
Past the benchmarks, Zhipu says it labored with safety groups in China to run its fashions towards actual codebases, figuring out 2,436 vulnerabilities throughout 269 open-source initiatives. The severity cut up is 107 crucial, 990 excessive, 1,286 medium and 53 low. The oldest flaw dates to 1981, and the typical vulnerability had been sitting in code for 26.6 years earlier than it was discovered.

One discrepancy is price carrying fastidiously. Zhipu’s abstract panel labels 1,097 findings as crucial and excessive, which matches the severity desk. The physique textual content of the identical launch describes these 1,097 as medium-to-high. A number of shops have reproduced the second model.
The depend additionally arrives after what Zhipu describes as professional evaluation, screening and deduplication, so the uncooked mannequin output isn’t what’s being reported. Of the two,436 findings, 53 have been publicly disclosed, and a couple of,383 stay underneath embargo. The discharge doesn’t say what number of had been beforehand unknown, and it doesn’t say what number of had been independently reproduced. These are the 2 figures that will flip a quantity declare right into a functionality declare.
What issues greater than the benchmark desk
Two issues within the launch have longer penalties than the CyberGym margin.
The primary is effectivity. Zhipu reviews GLM-5.3 reaching 31.4% on its inside coding benchmark at round 50,000 output tokens per job, towards Opus 4.8 at 29.5% utilizing 120,000. Barely higher work for lower than half the tokens is a value argument, and value determines whether or not safety groups exterior the most important budgets can run these instruments in any respect.
The second is distribution. Zhipu says the weights might be printed as soon as security analysis and hardening are completed. That has not occurred but, and till it does, the open-weights declare is a dedication somewhat than a reality. If it holds, a mannequin with documented vulnerability-discovery functionality turns into one thing any group can obtain and run domestically, together with in markets that may by no means have entry to an export-controlled American mannequin.
The weights are due on the finish of August.
See additionally: Anthropic walks into the White Home and Mythos is the rationale Washington let it in
Wish to study extra about AI and large knowledge from business leaders? Try AI & Big Data Expo going down in Amsterdam, California, and London. The excellent occasion is a part of TechEx and is co-located with different main know-how occasions together with the Cyber Security & Cloud Expo. Click on here for extra info.
AI Information is powered by TechForge Media. Discover different upcoming enterprise know-how occasions and webinars here.
