9 citations · 13 across the 5 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025★ 3 cited
Auditing language models for hidden objectives
Samuel Marks, Johannes Treutlein, Trenton Bricken +32
We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objecti…
cs.AI2025★ 1 cited
A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management
Simeon Campos, Henry Papadatos, Fabien Roger +3
The recent development of powerful AI systems has highlighted the need for robust risk management frameworks in the AI industry. Although companies have begun to implement safety f…