Moonshot AI’s Kimi K3 struggles to keep pace with leading US AI models in offensive cybersecurity tasks, a joint UK-US evaluation revealed on July 23. Despite its promising framework, Kimi K3 did not succeed in executing the most critical exploit outcomes in simulated tests.
Evaluation Results Highlight Gaps Against US Rivals
The UK’s AI Security Institute and the US Center of AI Standards and Innovation ran Kimi K3 through ExploitBench, a Carnegie Mellon benchmark assessing how an AI model advances software exploits. The test covered 41 vulnerabilities in Chrome’s V8 engine identified after 2023. Kimi K3 managed to complete only 32% of the tasks, while US leaders averaged a significantly higher 76.2%. In comparison, China’s GLM-5.2 scored 24%.
Most importantly, none of Kimi K3’s runs achieved arbitrary code execution, a critical phase granting attackers full control over a system. US models averaged this outcome in almost half of their attempts. Although Kimi K3 currently leads among open-weight models in this domain, its failure to reach this high-risk milestone signals a weak spot where real-world impact would be highest.
Performance in Network Breach Simulation
Another scenario, named “The Last Ones,” tested AI navigation through a complex corporate network consisting of about 20 hosts across four subnets. This 32-step cyberattack path usually requires about 20 hours for a human expert to complete fully.
Kimi K3 advanced to step 17 on average, whereas US competitors reached approximately step 28.5, and GLM-5.2 only managed 11 steps. Intriguingly, Kimi K3 successfully finished the entire sequence in one out of ten attempts, staying under the token limit set by evaluators, hinting at sporadic potential for full attacks.
Moonshot is set to release Kimi K3’s full weights publicly on July 27. This move will make the model’s capabilities accessible beyond any single platform. The joint report confirmed that the AI’s current safeguards did not sufficiently prevent it from assisting in offensive cyber activities, raising questions about future controls.



