3 papers
cs.CL2026
PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI
Snehasis Mukhopadhyay
Jailbreak attacks on multimodal AI systems remain underexplored, even though unsafe image generation can have more severe consequences than unsafe text and current defenses are rel…
cs.CL2026
DeceptGuard :A Constitutional Oversight Framework For Detecting Deception in LLM Agents
Snehasis Mukhopadhyay
Reliable detection of deceptive behavior in Large Language Model (LLM) agents is an essential prerequisite for safe deployment in high-stakes agentic contexts. Prior work on schemi…
cs.CL2025
AMBEDKAR-A Multi-level Bias Elimination through a Decoding Approach with Knowledge Augmentation for Robust Constitutional Alignment of Language Models
Snehasis Mukhopadhyay, Aryan Kasat, Shivam Dubey +5
Large Language Models (LLMs) can inadvertently reflect societal biases present in their training data, leading to harmful or prejudiced outputs. In the Indian context, our empirica…