5 citations · 11 across the 5 of their papers we have counts for
5 papers
Probing Warmth-Mediated Harm in Speech-Enabled LLMs for Mental-Health Conversations
Eugenia Kim, Bolor-Erdene Jagdagdorj, Dina Pekelis +2
Audio LLM benchmarks measure understanding and dialogue quality, not whether speech-enabled models respond with relational warmth when a vulnerable user discloses a mental-health c…
Lessons From Red Teaming 100 Generative AI Products
Blake Bullwinkel, Amanda Minnich, Shiven Chawla +23
In recent years, AI red teaming has emerged as a practice for probing the safety and security of generative AI systems. Due to the nascency of the field, there are many open questi…
PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
Gary D. Lopez Munoz, Amanda J. Minnich, Roman Lutz +17
Generative Artificial Intelligence (GenAI) is becoming ubiquitous in our daily lives. The increase in computational power and data availability has led to a proliferation of both s…
Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle
Emman Haider, Daniel Perez-Becker, Thomas Portet +28
Recent innovations in language model training have demonstrated that it is possible to create highly performant models that are small enough to run on a smartphone. As these models…
Online to Offline Crossover of White Supremacist Propaganda
Ahmad Diab, Bolor-Erdene Jagdagdorj, Lynnette Hui Xian Ng +2
White supremacist extremist groups are a significant domestic terror threat in many Western nations. These groups harness the Internet to spread their ideology via online platforms…