China's Z.ai claims new model outperforms Anthropic's Mythos 5 in cyber-defence tests
The model achieved 84.5% on CyberGym, slightly higher than Mythos 5's 83.8%. The results have not been independently verified. Z.ai's GLM-5.3 is being evaluated for delayed open release due to safety concerns.
Chinese AI startup Z.ai has announced that its new model, GLM-5.3, has performed nearly on par with Anthropic's Mythos 5 in cyber-defence tests. The model scored 84.5% on CyberGym, a benchmark that evaluates a model's ability to review code, identify security flaws, and confirm their validity. This performance slightly outperformed the 83.8% reported for Mythos 5, according to Z.ai's claims. However, these results have not been independently verified, leaving room for skepticism about their accuracy.
Z.ai's GLM-5.3 has demonstrated strong capabilities in identifying security flaws, but it has lagged behind Mythos 5 in converting these discoveries into working attacks. This is a standard part of cyber-defence testing, where models are evaluated not only on detection but also on their ability to simulate real-world threats. In a separate timed test, GLM-5.3 completed 105 attack-development tasks in two hours and 130 in six hours, showing a steady increase in performance over time.
The model's performance has sparked discussions about the safety and governance of AI systems. Z.ai's approach to delaying the open release of model weights is being viewed as a significant shift in the Chinese AI landscape. Gabriel Wagner, an AI governance researcher at Concordia AI, noted that this is the first time a Chinese lab has publicly justified such a delay with safety considerations. This move aligns with a broader trend of prioritizing responsible AI development.
The implications of Z.ai's approach could affect the broader AI industry, influencing how models are developed, tested, and released. The delayed open release of model weights may lead to increased costs for developers and users, as well as potential vendor lock-in. Additionally, the governance framework around AI models may become more complex, requiring greater oversight and regulation. Market reactions are likely to be mixed, with some viewing the approach as a necessary step toward safer AI and others seeing it as a barrier to innovation.
While Z.ai's GLM-5.3 has shown promising results in cyber-defence tests, the claims remain unverified, and the model's full capabilities are still under development. The company's strategy of balancing openness with safety considerations may set a precedent for other AI labs. However, the long-term success of this approach will depend on how effectively Z.ai can address the challenges of model transparency, security, and scalability.