collaborators

6 papers

cs.CL2026

Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models

Bocheng Chen, Han Zi, Roucheng Ou +5

In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural language attack surface unique t…

cs.CL2026

Learning to Diagnose and Correct Moral Errors: Beyond Shallow Heuristics in Moral Alignment

Bocheng Chen, Xi Chen, Han Zi +5

Existing approaches to moral value alignment are primarily set out to align LLMs' generation with the distributions of morally appropriate language, which has seen good progress. H…

cs.AI2026

Deactivating Refusal Triggers: Understanding and Mitigating Overrefusal in Safety Alignment

Zhiyu Xue, Zimo Qi, Guangliang Liu +2

Safety alignment aims to ensure that large language models (LLMs) refuse harmful requests by post-training on harmful queries paired with refusal answers. Although safety alignment…

cs.CL2026

Self-correction is Not An Innate Capability in Language Models

Guangliang Liu, Zimo Qi, Xitong Zhang +2

Although there has been growing interest in the self-correction capability of Large Language Models (LLMs), there are varying conclusions about its effectiveness. Prior research ha…

cs.CL2025

Discourse Heuristics For Paradoxically Moral Self-Correction

Guangliang Liu, Zimo Qi, Xitong Zhang +1

Moral self-correction has emerged as a promising approach for aligning the output of Large Language Models (LLMs) with human moral values. However, moral self-correction techniques…

cs.CL2025

Diagnosing Moral Reasoning Acquisition in Language Models: Pragmatics and Generalization

Guangliang Liu, Zimo Qi, Xitong Zhang +2

Ensuring that Large Language Models (LLMs) return just responses which adhere to societal values is crucial for their broader application. Prior research has shown that LLMs often…