3 papers
cs.AI2026
Robust Critics: Defending LLMs Against Multi-Turn Attacks
Roman Belaire, Arunesh Sinha, Pradeep Varakantham
When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the central challenges of LLM saf…
cs.LG2025
Automatic LLM Red Teaming
Roman Belaire, Arunesh Sinha, Pradeep Varakantham
Red teaming is critical for identifying vulnerabilities and building trust in current LLMs. However, current automated methods for Large Language Models (LLMs) rely on brittle prom…
cs.LG2025
On Minimizing Adversarial Counterfactual Error in Adversarial RL
Roman Belaire, Arunesh Sinha, Pradeep Varakantham
Deep Reinforcement Learning (DRL) policies are highly susceptible to adversarial noise in observations, which poses significant risks in safety-critical scenarios. The challenge in…