Mid-tier AI models now cross the hacking threshold
The models aren't catching up to frontier systems, they're getting good enough, and cheap enough, that running them repeatedly beats paying for the best.
TL;DR
XBOW research shows a class of proprietary and open-weight models, GLM-5.2, Grok 4.5, Opus 4.7, Muse Spark 1.1, now reliably complete "moderately complex" agentic hacking tasks that they couldn't handle six months ago. Their lower token costs let users run them many more times than a frontier model, sometimes leapfrogging the expensive option on net results. GPT 5.5 also posted one of XBOW's best-ever exploitation benchmark scores, and notably performed better without source-code access than GPT 5 did with it, working as an actual attacker would.
The finding that should shift how federal agencies think about AI threat models is the cost curve. Frontier models like Mythos and GPT 5.6 are more capable on individual tasks, but the token economics are brutal. Anthropic's own research this week found that a coordinating agent swarm of Mythos Preview agents found 266 vulnerabilities across 15 open-source projects, but burned 27 million tokens to get there. A team of agents working individually found 21 on 6.5 million tokens. Few organizations, and almost no individual threat actors, can underwrite that kind of research budget.
Mid-tier models invert the equation. They're weaker per inference but so much cheaper that running them repeatedly (giving them more time, more attempts, more parallel paths) often produces better outcomes than a single expensive frontier-model run. As XBOW's Albert Ziegler put it: "They come from behind and leapfrog the big frontier model."
Six months ago that strategy didn't work. Open-source models would get lost on long-horizon agentic tasks. Now they cross a threshold where they provide net value, and the economics take over from there.
GPT 5.5's performance on XBOW's exploitation benchmarks underscores the same dynamic from a different angle. Its "miss rate" (failure to spot a vulnerability) dropped to 10%, against GPT 5's 40%. More telling: it performed better without source-code access than GPT 5 did with it. "What translated into findings was the ability to reach and prove a vulnerability against the running system, not to infer it from a pattern in the source," the XBOW report said. That's the attacker's actual condition.
For federal CISOs and security teams, the practical implication is uncomfortable. Policies and risk assessments that focus on frontier-model access as the gating factor (who has API keys to the most advanced systems) miss the emerging threat surface. A competent adversary with a mid-tier open-weight model, modest compute, and patience can now do real damage. The models already exist; they're downloadable, not gated.
Published ·Deep Fathom