collaborators

5 papers

cs.SI2026

Understanding the Self-Reflection Mechanisms of LLMs through Biased Attitude Associations

Jingshen Zhang, Bo Wang, Boci Yang +4

While the emergent self-reflection capabilities of Large Language Models (LLMs) offer a promising paradigm for autonomous bias mitigation, their internal mechanics remain unclear,…

cs.SI2026

Modeling Implicit Conflict Monitoring Mechanisms against Stereotypes in LLMs

Jingshen Zhang, Bo Wang, Yanlin Fu +4

In this paper, we study an emergent self-debiasing mechanisms against stereotypical content in Large Language Models (LLMs). Unlike traditional safety mechanisms that are primarily…

cs.CL2024

Do LLMs Have the Generalization Ability in Conducting Causal Inference?

Chen Wang, Dongming Zhao, Bo Wang +2

In causal inference, generalization capability refers to the ability to conduct causal inference methods on new data to estimate the causal-effect between unknown phenomenon, which…

cs.CL2024

RoleBreak: Character Hallucination as a Jailbreak Attack in Role-Playing Systems

Yihong Tang, Bo Wang, Xu Wang +5

Role-playing systems powered by large language models (LLMs) have become increasingly influential in emotional communication applications. However, these systems are susceptible to…

cs.CL2024

MORPHEUS: Modeling Role from Personalized Dialogue History by Exploring and Utilizing Latent Space

Yihong Tang, Bo Wang, Dongming Zhao +4

Personalized Dialogue Generation (PDG) aims to create coherent responses according to roles or personas. Traditional PDG relies on external role data, which can be scarce and raise…