activity
20242026
collaborators

5 papers

cs.CY2026

Risk Reporting for Developers' Internal AI Model Use

Oscar Delaney, Sambhav Maheshwari, Joe O'Brien +2

Frontier AI companies first deploy their most advanced models internally, for weeks or months of safety testing, evaluation, and iteration, before a possible public release. For ex…

cs.CR2025

Asymmetry by Design: Boosting Cyber Defenders with Differential Access to AI

Shaun Ee, Chris Covino, Cara Labrador +3

As AI-enabled cyber capabilities become more advanced, we propose "differential access" as a strategy to tilt the cybersecurity balance toward defense by shaping access to these ca…

cs.CY2025

Expert Survey: AI Reliability & Security Research Priorities

Joe O'Brien, Jeremy Dolan, Jay Kim +5

Our survey of 53 specialists across 105 AI reliability and security research areas identifies the most promising research prospects to guide strategic AI R&D investment. As compani…

cs.LG2025

Cost and Reward Infused Metric Elicitation

Chethan Bhateja, Joseph O'Brien, Afnaan Hashmi +1

In machine learning, metric elicitation refers to the selection of performance metrics that best reflect an individual's implicit preferences for a given application. Currently, me…

cs.CY2024

Coordinated Disclosure of Dual-Use Capabilities: An Early Warning System for Advanced AI

Joe O'Brien, Shaun Ee, Jam Kraprayoon +3

Advanced AI systems may be developed which exhibit capabilities that present significant risks to public safety or security. They may also exhibit capabilities that may be applied…