Seventy-three percent. That's the share of AI-augmented teams that solved at least one challenge in a recent 72-hour NeuroGrid cybersecurity test, a margin that shows agents already outperform most human competitors. The jump in baseline capability arrived as Anthropic said it would limit access to its new Claude Mythos model to a small set of trusted organizations, and OpenAI said it would share the most capable versions of its systems only with a narrowed partner set. The scramble to control access matters because the same competitions that highlight agent speed also reveal persistent gaps on the hardest tasks and fresh risks for training and operations.
Anthropic and OpenAI tightened distribution of their top models amid a string of public exercises that placed autonomous AI agents into both offensive and defensive roles in live cyber operations. Anthropic announced last month that it would restrict Claude Mythos to a small number of trusted organizations. OpenAI said it would similarly narrow partners for the most advanced versions of its systems. Those moves came as event organizers and researchers published results showing agents can multiply human throughput on routine work, while falling short against elite human practitioners on the most complex engagements.
Where agents win, and where they stall
Large-scale benchmark work on the Hack The Box platform under the NeuroGrid label supplies the clearest statistical snapshot. NeuroGrid ran a 72-hour event that registered 1,337 human-only teams and 156 AI-agent teams. Organizers analyzed 958 human teams and 120 AI-agent teams and found roughly 73 percent of AI-augmented teams completed at least one challenge versus 46 percent for human-only teams. The advantage was strongest on lower-complexity tasks and on medium-difficulty problems where mid-career analysts typically operate. As participant skill rose, the edge narrowed. NeuroGrid reported the AI solve-rate advantage fell from about 3.2x across the field to 1.69x among the top 5 percent of teams.
Complementary work from Palisade Research produced similar but distinct results. In one Capture The Flag style run, four of seven AI agents solved 19 out of 20 tasks during a 48-hour period, and the top AI finished in the top five percent overall. In a larger follow-up, an AI called CAI solved 20 out of 62 tasks and placed around 859th among roughly 18,000 participants. That ranking put CAI in the top ten percent of all teams and ahead of about 90 percent of competing humans on those metrics. Palisade’s testing made a technical limitation clear. Tasks that required interaction with external machines or environments were markedly harder for agents designed to execute locally, which reduced effectiveness on the most complex, interactive engagement types.
Speed and workflow differences matter as well. NeuroGrid found AI-augmented teams were marginally slower on average, but top AI teams completed elite-level challenges several times faster than their human counterparts. Practitioners described the practical effect as force multiplication rather than replacement. Red-team operator Dan Borges summed this up. “They help me do things in parallel. I can go fast, and I can go wide,” he said. Alex Levinson, a red-team leader at the collegiate exercise, explained how scoring works: stealing data or gaining access imposed direct penalties on defending blue teams, which keeps the contests close to real-world incentives.
Operational fragility and workforce ripple effects
Even where agents score well, they behave unpredictably in live settings. On-site reporting from a Cosmopolitan suite exercise recorded at least one incident in which a bot "took an unexpected turn," a phrase that reflects the kinds of brittle failures organizers are still seeing in unsupervised operations. Agents can make mistakes that human supervisors must catch.
That fragility matters when teams use AI.
Organizers and analysts are also worried about how automation reshapes training pipelines. NeuroGrid and other observers warned that routine tasks are the classroom for junior analysts. Entry-level work that agents can handle today is where early-career staff learn the judgment, verification, and skeptical instincts required for senior roles. If organizations over-deploy agents for training-era work, they risk hollowing out the next generation of human expertise.
At the same time, the data show elite human teams still hold important advantages. Creative, open-ended tasks such as reverse engineering and exploit development narrow the gap between humans and machines, and at the highest level there's clear parity or superior human performance. That suggests a persistent role for seasoned practitioners in complex incident response and adversary simulation. The operational read is straightforward. Use agents to accelerate repetitive and mid-skill work, keep humans engaged on the hard problems, and build oversight to catch the bots’ mistakes.
The evidence isn't uniform. NeuroGrid’s population-scale statistics drive many of the workforce and policy takeaways, while Palisade’s tournaments emphasize different task sets and metrics.
The collegiate competition coverage supplies qualitative color, showing agents can be pasted into blue teams while red teams use AI to augment attacks. Together, the datasets describe a hybrid future where agents are powerful tools, not replacements.
The twin announcements by Anthropic and OpenAI change the commercial framing. By limiting access to their top models, both companies are shifting high-capability systems toward a partner-first distribution model, at least for now. That narrows who will be able to deploy agent stacks at scale, and keeps pricing and broader availability unspecified while the community debates safety and operational best practices.
Related Articles
- Anthropic offers free Claude courses while reshaping billing and adding SpaceX compute
- OpenAI model sparks compute debate, shakes cloud deals
- Royal Pop: 8 Audemars Piguet x Swatch Pocket Watches, $400
Anthropic has limited access to Claude Mythos to a small set of trusted organizations, and OpenAI said it will share equivalent high-capability systems only with a narrowed partner set, with wider availability and pricing left unspecified.
This article was created with AI assistance.