Showing cs.CYShow all
2 papers · 1 filter
cs.CY2025
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
Ann-Kathrin Dombrowski, Dillon Bowen, Adam Gleave +1
Open-weight large language models (LLMs) unlock huge benefits in innovation, personalization, privacy, and democratization. However, their core advantage - modifiability - opens th…
cs.CY2025
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
Dillon Bowen, Ann-Kathrin Dombrowski, Adam Gleave +1
The rapid advancement of AI systems has raised widespread concerns about potential harms of frontier AI systems and the need for responsible evaluation and oversight. In this posit…