The cyber gap is shrinking
The UK government's AI Security Institute has analyzed the delta in cybersecurity capabilities between powerful proprietary models and open-weight models. The results show that this year, the gap has shrunk.
"This is our first public analysis of how far leading open weight models trail the closed cyber frontier," AISI writes. "Recent open models GLM-5.2 and DeepSeek V4-Pro perform similarly to frontier closed models released 4 to 7 months before them, a narrower gap than the 6 to 10 months we measured through most of 2025."
On a set of 70 evaluations for specific, narrow cyber capabilities, GLM-5.2 is closest to Claude Opus 4.6, which was released 4.3 months earlier. DeepSeek-V4-Pro sits somewhere between Claude Opus 4.5 and GPT-5, released in November and August 2025, respectively. AISI intends to test Kimi K3 on the same basis once its weights are publicly released.
The gap lengthens a bit for long-horizon cyber ranges, which test how well models can chain capabilities to complete a full hacking operation. On a cyberrange called The Last Ones, GLM-5.2 reaches as far as Opus 4.5, a model released less than 7 months before it, while DeepSeek V4-Pro falls below Sonnet 4.5.
Kimi K3 enters the chat
In the last couple of years, Chinese firms have begun to outcompete Western actors at building and deploying open-weight models, and are now starting to close the gap on frontier models as well. The latest and best example is Kimi K3, a 2.8 trillion parameter model. Kimi has exceptionally strong scores on all the tasks that major proprietary models benchmark on and typically matches or trails Claude Fable 5 and GPT 5.6 Sol.
"While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models," Kimi writes. The weights will be made available in the coming weeks along with a research paper about the model.
Demis Hassabis proposes a Standards Body
DeepMind founder Demis Hassabis has laid out a policy prescription for AGI. His basic idea is that the US government should develop a framework for testing frontier AI systems for new capabilities, and should do this via a Standards Body modeled on a federally overseen public-private partnership or self-regulatory organization, much like the Financial Industry Regulatory Authority (FINRA).
"This US-initiated effort would provide a strong starting point for creating shared international standards on Frontier AI," he says. The Standards Body would be responsible for developing assessment protocols and working with appropriate federal agencies and the US National Labs to conduct testing in areas relevant to national security. This infrastructure would help define what makes a model a "Frontier Model," and labs developing those models would be encouraged to adopt best practices in areas like publishing details about their systems, investing in cybersecurity, and personnel vetting.
"Initially, Frontier Labs would voluntarily share models with the Standards Body for review up to 30 days before release. Once the assessment protocol is shown to be effective and robust, formalization could quickly follow," he writes.
The proposal is a moderate one: voluntary, US-led, modeled on a regulator that has worked in finance. The harder questions, like whether frontier labs should be required to share, and what the enforcement mechanism looks like, are deliberately left for later. The interesting question is whether other labs will publicly support it before the open-weight releases they can no longer control force the issue.
The structural shift behind the news
Each of the three stories in Import AI 465 is interesting on its own. Together, they tell a more important story about where the AI field is in 2026. The cyber gap is narrowing because open-weight model developers are no longer years behind the frontier on training compute, data, and engineering talent. Kimi K3 is the proof point that a Chinese lab can build a model that competes with the best closed Western systems on the benchmarks that matter. The Hassabis proposal is the recognition that voluntary standards are no longer enough, and that some form of structured public oversight is coming whether the labs want it or not.
The interesting question is not whether any of these three threads will continue. All of them will. The interesting question is what the world looks like when the cyber gap closes entirely, when multiple labs can ship frontier-tier models in the same week, and when a US-led Standards Body is actually evaluating pre-release frontier systems. None of those is far off.
The next Import AI will probably cover more of the same, and the one after that. The pace is the story, and the story is the pace.