170 citations · 289 across the 3 of their papers we have counts for
3 papers
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Deep Ganguli, Liane Lovitt, Jackson Kernion +33
We describe our early efforts to red team language models in order to simultaneously discover, measure, and attempt to reduce their potentially harmful outputs. We make three main…
Language Models (Mostly) Know What They Know
Saurav Kadavath, Tom Conerly, Amanda Askell +33
We study whether language models can evaluate the validity of their own claims and predict which questions they will be able to answer correctly. We first show that larger models a…
Long-Term Benefits of Network Boosters for Renewables Integration and Corrective Grid Security
Amin Shokri Gazafroudi, Elisabeth Zeyen, Martha Frysztacki +2
The preventative strategies for network security dominant in European networks mean that network capacity is kept free in case a line fails. If instead fast corrective action…