Preparing for Offensive Cyber Agents: Piloting AI-Enabled Penetration Testing and Red Teaming

This report was coauthored by Jam Kraprayoon and Matthew Mittelsteadt.

Offensive cyber agents that can autonomously execute cyberattacks are an emerging national security threat. Recent incidents disclosed by major AI developers and evaluators show models executing cyberattacks with increasing autonomy. In August 2026, OpenAI reported that an unreleased model may have reached the company’s internal definition of “critical cyber capability threshold,” which covers models able to devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level desired goal. Threat actors could leverage these models to develop and deploy increasingly autonomous offensive cyber agents, combining the speed and scale of machines with the adaptability of human operators.

The current cyber landscape would be highly vulnerable to these agents. Agents could systematically exploit known weaknesses at scale, such as unpatched and misconfigured systems. These vulnerabilities are endemic in large enterprises and among critical but under-resourced defenders, such as public utilities and hospitals. Against better-defended targets, they could adapt and persist. Less constrained by the physical limits of human operators and rule-based automation, these agents could launch more frequent and successful attacks, coordinated strikes on critical infrastructure, or highly adaptive worms. These same weaknesses could also enable loss-of-control scenarios, where an agent acting against an operator’s intent attacks systems on its own.

As threat actors adopt increasingly capable offensive agents, U.S. cyber policy needs to prioritize hardening the cyber landscape. Penetration testing (or “pentesting”) and red teaming complement broader efforts to secure software, manage vulnerabilities, and improve detection and response. They are proven methods for prioritizing remediation and hardening defenses, finding the attack paths threat actors could take to mission-critical systems and testing a defender’s ability to respond. But these engagements are labor and time intensive, and many defenders cannot afford regular testing. Federal services can fill this gap for critical but under-resourced defenders, but face the same resource and labor constraints. These engagements also emulate human-led attacks, not the speed and novelty of agentic ones.

AI-enabled pentesting and red teaming, which uses AI agents and tools to conduct these assessments, is an emerging AI application that could address many of these challenges and help defenders keep pace with offensive adoption. These systems could act as a force multiplier for human teams, helping identify vulnerable systems and attack paths faster, driving down the cost of assessments, enabling more frequent and continuous testing, and emulating the speed and novelty of agentic attacks that human testers cannot replicate.

AI models, agents, and tooling for pentesting and red teaming are developing rapidly, but a development-to-deployment gap will keep defenders from taking advantage of that progress. Emulating attacks on a defender’s own systems and networks with higher-autonomy agents is high-risk activity that could disrupt operations. Many defenders will hesitate to deploy these agents without the appropriate safeguards and governance in place.

Federal pilots can accelerate federal and non-federal deployment by testing these systems in real-world conditions and identifying the technical and governance practices needed to operationalize higher-autonomy agents at scale. Human operators would experiment with and evaluate these systems in controlled environments and non-critical networks, surfacing and solving the problems that must be addressed for widespread deployment. Congress should direct the Cybersecurity and Infrastructure Security Agency (CISA) and National Security Agency (NSA), in collaboration with other relevant federal agencies, to conduct operational pilots working toward the following four goals:

Figure 1. Summary of Federal Pilot Goals and Outcomes

Endnote

  1.  A harness (or scaffolding) is the software wrapped around a model that turns it into an agent: the tools, memory, and control logic that let it take actions, observe results, and plan multiple steps. The same model can perform far better or worse depending on the harness built around it.

Next
Next

Differential Automation: Steering Automated Research and Development Toward Safety and Security