5 papers
Risk Reporting for Developers' Internal AI Model Use
Oscar Delaney, Sambhav Maheshwari, Joe O'Brien +2
Frontier AI companies first deploy their most advanced models internally, for weeks or months of safety testing, evaluation, and iteration, before a possible public release. For ex…
Asymmetry by Design: Boosting Cyber Defenders with Differential Access to AI
Shaun Ee, Chris Covino, Cara Labrador +3
As AI-enabled cyber capabilities become more advanced, we propose "differential access" as a strategy to tilt the cybersecurity balance toward defense by shaping access to these ca…
Expert Survey: AI Reliability & Security Research Priorities
Joe O'Brien, Jeremy Dolan, Jay Kim +5
Our survey of 53 specialists across 105 AI reliability and security research areas identifies the most promising research prospects to guide strategic AI R&D investment. As compani…
Cost and Reward Infused Metric Elicitation
Chethan Bhateja, Joseph O'Brien, Afnaan Hashmi +1
In machine learning, metric elicitation refers to the selection of performance metrics that best reflect an individual's implicit preferences for a given application. Currently, me…
Coordinated Disclosure of Dual-Use Capabilities: An Early Warning System for Advanced AI
Joe O'Brien, Shaun Ee, Jam Kraprayoon +3
Advanced AI systems may be developed which exhibit capabilities that present significant risks to public safety or security. They may also exhibit capabilities that may be applied…