4 papers
Understanding the Self-Reflection Mechanisms of LLMs through Biased Attitude Associations
Jingshen Zhang, Bo Wang, Boci Yang +4
While the emergent self-reflection capabilities of Large Language Models (LLMs) offer a promising paradigm for autonomous bias mitigation, their internal mechanics remain unclear,…
Modeling Implicit Conflict Monitoring Mechanisms against Stereotypes in LLMs
Jingshen Zhang, Bo Wang, Yanlin Fu +4
In this paper, we study an emergent self-debiasing mechanisms against stereotypical content in Large Language Models (LLMs). Unlike traditional safety mechanisms that are primarily…
Do LLMs Have the Generalization Ability in Conducting Causal Inference?
Chen Wang, Dongming Zhao, Bo Wang +2
In causal inference, generalization capability refers to the ability to conduct causal inference methods on new data to estimate the causal-effect between unknown phenomenon, which…
Think Twice: A Human-like Two-stage Conversational Agent for Emotional Response Generation
Yushan Qian, Bo Wang, Shangzhao Ma +5
Towards human-like dialogue systems, current emotional dialogue approaches jointly model emotion and semantics with a unified neural network. This strategy tends to generate safe r…