2 citations · 2 across the 2 of their papers we have counts for
5 papers
RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
Ali Asad, Stephen Obadinma, Radin Shayanfar +1
We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate their unsafe behaviours. AI safe…
Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception
Ali Asad, Stephen Obadinma, Anshul Pattoo +2
Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally induced goal. Yet it remains unclear how con…
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
Stephen Obadinma, Xiaodan Zhu
Robust verbal confidence generated by large language models (LLMs) is crucial for the deployment of LLMs to help ensure transparency, trust, and safety in many applications, includ…
FAIIR: Building Toward A Conversational AI Agent Assistant for Youth Mental Health Service Provision
Stephen Obadinma, Alia Lachana, Maia Norman +7
The world's healthcare systems and mental health agencies face both a growing demand for youth mental health services, alongside a simultaneous challenge of limited resources. Here…
Calibration Attacks: A Comprehensive Study of Adversarial Attacks on Model Confidence
Stephen Obadinma, Xiaodan Zhu, Hongyu Guo
In this work, we highlight and perform a comprehensive study on calibration attacks, a form of adversarial attacks that aim to trap victim models to be heavily miscalibrated withou…