Benchmark shows Chinese AI closing gap in bug hunting
A Chinese AI model is matching Anthropic’s Mythos in some bug-finding tests, raising policy and security questions as open-weight tools spread.
Jason Kwon ·

A Chinese AI model has reached parity with Anthropic’s Mythos in select cybersecurity evaluations, researchers say, sharpening competition and policy debate.
The results center on Zhipu AI’s newly released GLM-5.2, which security researchers and benchmarking groups say can perform at the level of top U.S. systems when tasked with identifying software vulnerabilities. While it still trails leading American models on broader capabilities, the bug-hunting progress is drawing attention because it directly affects how quickly defenses—and attacks—can scale.
Bug-finding benchmarks show a tighter U.S.-China race
Security researchers said GLM-5.2 can match the latest U.S. models in scenarios focused on locating security flaws in code. Lior Div, chief executive of cybersecurity firm 7AI, said China is steadily reducing the distance to the most advanced American systems.
Semgrep, a company that evaluates code-security tools, reported that GLM-5.2 outperformed Anthropic’s Claude Opus 4.8 in some benchmark tests. Opus 4.8 was released in May, and the comparisons suggest rapid iteration on the Chinese side over a short timeframe.
Researchers also said that when the models receive additional instructions, both Opus 4.8 and GLM-5.2 can reach Mythos-level performance in bug discovery. That matters because vulnerability research is one of the clearest near-term paths for AI to create real-world security impact—both by helping developers fix issues and by enabling adversaries to find weaknesses faster.
Open-weight design boosts adoption—and raises misuse risks
Unlike Anthropic’s or OpenAI’s most capable offerings, GLM-5.2 is described as open-weight, meaning the model can be downloaded, run on privately controlled hardware, and modified. This design appeals to organizations that want system control, predictable deployment, and reduced reliance on external cloud providers.
However, the same attributes also lower barriers for criminal misuse. A model that can be operated outside centralized monitoring can be used discreetly to automate vulnerability discovery, craft exploit paths, or scale reconnaissance workflows.
Concerns about a rapid increase in exploitable weaknesses have been framed by some researchers as the risk of a “bugmageddon,” a scenario where the rate of newly discovered vulnerabilities overwhelms patching capacity. Stronger AI for bug finding increases pressure on software teams and security vendors to shorten remediation cycles and improve secure-by-design practices.
Market distribution and policy stakes rise in the U.S.
Adoption indicators suggest GLM-5.2 is already spreading quickly. OpenRouter, which aggregates access to more than 400 AI models, lists GLM-5.2 among its 10 most-used models.
Growing usage is also being fueled by economics, as businesses look for ways to contain AI spending. Researchers and industry watchers say cost sensitivity is accelerating interest in Chinese systems, and large technology firms—including Microsoft—are evaluating how Chinese models could be offered through their platforms.
New tooling is emerging in parallel. On Wednesday, Chinese cybersecurity company 360 Security Technology introduced a bug-finding product called Tulongfeng, saying its performance is comparable to Mythos in vulnerability discovery.
National-security officials and corporate leaders have been alarmed by the implications of broadly available, high-performing security-focused models. The developments also arrive as the White House is revisiting U.S. AI policy, increasing scrutiny on how model access, export controls, platform distribution, and security testing standards may evolve.
Next steps to watch include whether additional independent benchmarks confirm the most competitive results, whether major platforms expand distribution of Chinese models, and how U.S. policy updates address open-weight systems that can be deployed without centralized oversight.