20 citations · 22 across the 4 of their papers we have counts for
Showing cs.CRShow all
2 papers · 1 filter
cs.CR2025
Prmpt: Sanitizing Sensitive Prompts for LLMs
Amrita Roy Chowdhury, David Glukhov, Divyam Anshumaan +4
The rise of large language models (LLMs) has introduced new privacy challenges, particularly during inference where sensitive information in prompts may be exposed to proprietary L…
cs.CR2024
Breach By A Thousand Leaks: Unsafe Information Leakage in `Safe' AI Responses
David Glukhov, Ziwen Han, Ilia Shumailov +2
Vulnerability of Frontier language models to misuse and jailbreaks has prompted the development of safety measures like filters and alignment training in an effort to ensure safety…