2 citations · 2 across the 2 of their papers we have counts for
3 papers
From Risk Avoidance to User Empowerment in AI Mental Health Crisis Support
Benjamin Kaveladze, Arka Ghosh, Leah Ajmani +11
People experiencing mental health crises frequently turn to open-ended generative AI (GenAI) chatbots for support. However, rather than providing immediate assistance, some GenAI c…
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
Blake Bullwinkel, Mark Russinovich, Ahmed Salem +8
Recent research has demonstrated that state-of-the-art LLMs and defenses remain susceptible to multi-turn jailbreak attacks. These attacks require only closed-box model access and…
Lessons From Red Teaming 100 Generative AI Products
Blake Bullwinkel, Amanda Minnich, Shiven Chawla +23
In recent years, AI red teaming has emerged as a practice for probing the safety and security of generative AI systems. Due to the nascency of the field, there are many open questi…