activity
20242026
most citedRedDebate: Safer Responses Through Multi-Agent Red Teaming Debates

2 citations · 2 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CL20262 cited

RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates

Ali Asad, Stephen Obadinma, Radin Shayanfar +1

We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate their unsafe behaviours. AI safe…

cs.CL2026

Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception

Ali Asad, Stephen Obadinma, Anshul Pattoo +2

Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally induced goal. Yet it remains unclear how con…

cs.CL2025

On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks

Stephen Obadinma, Xiaodan Zhu

Robust verbal confidence generated by large language models (LLMs) is crucial for the deployment of LLMs to help ensure transparency, trust, and safety in many applications, includ…

cs.AI2025

FAIIR: Building Toward A Conversational AI Agent Assistant for Youth Mental Health Service Provision

Stephen Obadinma, Alia Lachana, Maia Norman +7

The world's healthcare systems and mental health agencies face both a growing demand for youth mental health services, alongside a simultaneous challenge of limited resources. Here…

cs.LG2024

Calibration Attacks: A Comprehensive Study of Adversarial Attacks on Model Confidence

Stephen Obadinma, Xiaodan Zhu, Hongyu Guo

In this work, we highlight and perform a comprehensive study on calibration attacks, a form of adversarial attacks that aim to trap victim models to be heavily miscalibrated withou…