10 citations · 10 across the 1 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2024
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
Cameron Tice, Philipp Alexander Kreer, Nathan Helm-Burger +7
Capability evaluations play a crucial role in assessing and regulating frontier AI systems. The effectiveness of these evaluations faces a significant challenge: strategic underper…
cs.AI2023★ 10 cited
Will releasing the weights of future large language models grant widespread access to pandemic agents?
Anjali Gopal, Nathan Helm-Burger, Lennart Justen +6
Large language models can benefit research and human understanding by providing tutorials that draw on expertise from many different fields. A properly safeguarded model will refus…