2 citations · 3 across the 3 of their papers we have counts for
4 papers · 1 filter
Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception
Ali Asad, Stephen Obadinma, Anshul Pattoo +2
Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally induced goal. Yet it remains unclear how con…
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
Stephen Obadinma, Xiaodan Zhu
Robust verbal confidence generated by large language models (LLMs) is crucial for the deployment of LLMs to help ensure transparency, trust, and safety in many applications, includ…
RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
Ali Asad, Stephen Obadinma, Radin Shayanfar +1
We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate their unsafe behaviours. AI safe…
SemEval-2020 Task 5: Counterfactual Recognition
Xiaoyu Yang, Stephen Obadinma, Huasha Zhao +3
We present a counterfactual recognition (CR) task, the shared Task 5 of SemEval-2020. Counterfactuals describe potential outcomes (consequents) produced by actions or circumstances…