ARTFEED — Contemporary Art Intelligence

Moonshot AI's Kimi K3 trails US models in cyberattack capability, UK-US study finds

ai-technology · 2026-07-24

A study from the UK Artificial Intelligence Security Institute combined with the US Centre for AI Standards and Innovation has revealed that China’s Moonshot AI’s Kimi K3 language model is underperforming in cybersecurity evaluations. Achieving a score of only 32.2% on the ExploitBench benchmark, Kimi K3 falls short of American counterparts that average 76.2%. Although it surpassed Zhipu AI’s GLM-5.2 at 24.4%, Kimi K3 failed to complete any of the 41 tasks requiring code execution, unlike US models, which successfully executed code in 20 instances. The report, published Thursday, raises alarms about the advancement of Chinese open-source AI technology.

Key facts

  • Kimi K3 scored 32.2% on ExploitBench
  • Top US models averaged 76.2%
  • Kimi K3 failed arbitrary code execution on all 41 tasks
  • US models achieved arbitrary code execution on 20 tasks
  • Zhipu AI's GLM-5.2 scored 24.4%
  • Study by UK AISI and US CAISI
  • Report published on Thursday
  • Kimi K3 is considered China's most powerful LLM

Entities

Institutions

  • Moonshot AI
  • UK Artificial Intelligence Security Institute (AISI)
  • US Centre for AI Standards and Innovation (CAISI)
  • UK Department for Science, Innovation and Technology
  • US Department of Commerce
  • National Institute of Standards and Technology
  • Zhipu AI

Locations

  • China
  • United Kingdom
  • United States
  • Beijing

Sources