4 papers
Attention Is Where You Attack
Aviral Srivastava, Sourav Panda
Safety-aligned large language models rely on RLHF and instruction tuning to refuse harmful requests, yet the internal mechanisms implementing safety behavior remain poorly understo…
Signed, Sealed,... Confused: Exploring the Understandability and Severity of Policy Documents
Shikha Soneji, Sourav Panda, Sameer Neve +1
In general, Terms of Service (ToS) and other policy documents are verbose and full of legal jargon, which poses challenges for users to understand. To improve user accessibility an…
A Formal Framework for Assessing and Mitigating Emergent Security Risks in Generative AI Models: Bridging Theory and Dynamic Risk Mitigation
Aviral Srivastava, Sourav Panda
As generative AI systems, including large language models (LLMs) and diffusion models, advance rapidly, their growing adoption has led to new and complex security risks often overl…
How to Measure Human-AI Prediction Accuracy in Explainable AI Systems
Sujay Koujalgi, Andrew Anderson, Iyadunni Adenuga +8
Assessing an AI system's behavior-particularly in Explainable AI Systems-is sometimes done empirically, by measuring people's abilities to predict the agent's next move-but how to…