activity
20232026
most citedFairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language Models

3 citations · 4 across the 13 of their papers we have counts for

collaborators
Showing 2025Show all

5 papers · 1 filter

cs.CY2025

Uncovering Strategic Egoism Behaviors in Large Language Models

Yaoyuan Zhang, Aishan Liu, Zonghao Ying +4

Large language models (LLMs) face growing trustworthiness concerns (\eg, deception), which hinder their safe deployment in high-stakes decision-making scenarios. In this paper, we…

cs.CL2025

Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing

Yisong Xiao, Aishan Liu, Siyuan Liang +3

Large Language Models (LLMs) have demonstrated impressive performance across various tasks, yet they remain vulnerable to generating toxic content, necessitating detoxification str…

cs.CR2025

PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking

Quanchen Zou, Zonghao Ying, Moyang Chen +7

The increasing sophistication of large vision-language models (LVLMs) has been accompanied by advances in safety alignment mechanisms designed to prevent harmful content generation…

cs.SE2025★ 3 cited

Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language Models

Yisong Xiao, Aishan Liu, Siyuan Liang +2

LLMs have demonstrated remarkable performance across diverse applications, yet they inadvertently absorb spurious correlations from training data, leading to stereotype association…

cs.CL2025

Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models

Zonghao Ying, Deyue Zhang, Zonglei Jing +7

Multi-turn jailbreak attacks simulate real-world human interactions by engaging large language models (LLMs) in iterative dialogues, exposing critical safety vulnerabilities. Howev…