Theo Bearman Theo Bearman

The OpenAI/Hugging Face Incident: Challenges in Controlling and Containing Cyber-Capable AI Systems

The OpenAI/Hugging Face incident is the first known instance of an AI system acting outside its developer’s intentions to autonomously identify a target and execute an attack end-to-end. Policymakers should consider it a warning shot that has left critical questions unanswered. In the incident’s wake, Congress should seek industry-wide responses to questions outlined in this memo, which also sets out actions that policymakers can take today to mitigate risks.

Read More
Oscar Delaney Oscar Delaney

Risk Reporting for Developers’ Internal AI Model Use

Frontier AI companies run their most capable models internally for weeks before public release. This report offers a harmonized reporting standard for internal use risks across SB 53, RAISE, and the EU Code of Practice.

Read More
Research Report Shaun Ee Research Report Shaun Ee

Asymmetry by Design: Boosting Cyber Defenders with Differential Access to AI

“Differential access” is a strategy to tilt the cybersecurity balance toward defense by shaping access to advanced AI-powered cyber capabilities. We introduce three possible approaches, Promote Access, Manage Access, and Deny by Default, with one constant across all approaches — even in the most restrictive scenarios, developers should aim to advantage cyber defenders.

Read More
Research Report Joe O'Brien Research Report Joe O'Brien

Coordinated Disclosure of Dual-Use Capabilities: An Early Warning System for Advanced AI

Future AI systems may be capable of enabling offensive cyber operations, lowering the barrier to entry for designing and synthesizing bioweapons, and other high-consequence dual-use applications. If and when these capabilities are discovered, who should know first, and how? We describe a process for information-sharing on dual-use capabilities and make recommendations for governments and industry to develop this process.

Read More