7 citations · 7 across the 2 of their papers we have counts for
2 papers
cs.CL2025★ 7 cited
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Mrinank Sharma, Meg Tong, Jesse Mu +40
Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes…
cs.LG2023
Adversarial Predictions of Data Distributions Across Federated Internet-of-Things Devices
Samir Rajani, Dario Dematties, Nathaniel Hudson +4
Federated learning (FL) is increasingly becoming the default approach for training machine learning models across decentralized Internet-of-Things (IoT) devices. A key advantage of…