2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CL2026★ 2 cited
RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
Ali Asad, Stephen Obadinma, Radin Shayanfar +1
We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate their unsafe behaviours. AI safe…
cs.CL2026
Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception
Ali Asad, Stephen Obadinma, Anshul Pattoo +2
Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally induced goal. Yet it remains unclear how con…