2 papers
cs.CR2026
Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs
Anirudh Sekar, Mrinal Agarwal, Rachel Sharma +4
Prompt injection attacks have become an increasing vulnerability for LLM applications, where adversarial prompts exploit indirect input channels such as emails or user-generated co…
cs.MA2025
WOLF: Werewolf-based Observations for LLM Deception and Falsehoods
Mrinal Agarwal, Saad Rana, Theo Sundoro +5
Deception is a fundamental challenge for multi-agent reasoning: effective systems must strategically conceal information while detecting misleading behavior in others. Yet most eva…