7 citations · 8 across the 3 of their papers we have counts for
3 papers
cs.LG2025
Forecasting Rare Language Model Behaviors
Erik Jones, Meg Tong, Jesse Mu +7
Standard language model evaluations can fail to capture risks that emerge only at deployment scale. For example, a model may produce safe responses during a small-scale beta test,…
cs.CL2025★ 7 cited
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Mrinank Sharma, Meg Tong, Jesse Mu +40
Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes…
stat.ML2019★ 1 cited
Differentially Private Federated Variational Inference
Mrinank Sharma, Michael Hutchinson, Siddharth Swaroop +2
In many real-world applications of machine learning, data are distributed across many clients and cannot leave the devices they are stored on. Furthermore, each client's data, comp…