4 papers
Context Inference Attacks Without Jailbreaks
Prince Jha, Samuele Poppi, Nils Lukas
Agentic AI systems are increasingly deployed to process sensitive data at inference time, such as healthcare records or financial documents assembled into a hidden \emph{context} b…
Boundary-targeted Membership Inference Attacks on Safety Classifiers
Anthony Hughes, Alexander Goldberg, Prince Jha +3
Safety classifiers are essential safeguards within generative AI systems, filtering harmful content or identifying at-risk users when interacting with large language models. Despit…
On the Cultural Anachronism and Temporal Reasoning in Vision Language Models
Mukul Ranjan, Prince Jha, Khushboo Kumari +1
Vision-Language Models (VLMs) are increasingly applied to cultural heritage materials, from digital archives to educational platforms. This work identifies a fundamental issue in h…
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
Prince Jha, Raghav Jain, Konika Mandal +3
In the digital world, memes present a unique challenge for content moderation due to their potential to spread harmful content. Although detection methods have improved, proactive…