2 citations · 2 across the 2 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026★ 2 cited
RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
Ali Asad, Stephen Obadinma, Radin Shayanfar +1
We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate their unsafe behaviours. AI safe…
cs.CL2026
Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception
Ali Asad, Stephen Obadinma, Anshul Pattoo +2
Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally induced goal. Yet it remains unclear how con…
cs.CL2025
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
Stephen Obadinma, Xiaodan Zhu
Robust verbal confidence generated by large language models (LLMs) is crucial for the deployment of LLMs to help ensure transparency, trust, and safety in many applications, includ…