4 papers
Attention Is Where You Attack
Aviral Srivastava, Sourav Panda
Safety-aligned large language models rely on RLHF and instruction tuning to refuse harmful requests, yet the internal mechanisms implementing safety behavior remain poorly understo…
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking
Harsha Vardhan Khurdula, Basem Rizk, Indus Khaitan +3
Current benchmarks for evaluating Vision Language Models (VLMs) often fall short in thoroughly assessing model abilities to understand and process complex visual and textual conten…
A Formal Framework for Assessing and Mitigating Emergent Security Risks in Generative AI Models: Bridging Theory and Dynamic Risk Mitigation
Aviral Srivastava, Sourav Panda
As generative AI systems, including large language models (LLMs) and diffusion models, advance rapidly, their growing adoption has led to new and complex security risks often overl…
Discovering an invisible Z' at the muon collider
Anjan Kumar Barik, Santosh Kumar Rai, Aviral Srivastava
We show in this letter how a heavy invisible gauge boson that will practically be out of reach of the Large Hadron Collider (LHC), can be discovered at th…