activity
20242026
collaborators

7 papers

cs.CL2026

Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models

Bocheng Chen, Han Zi, Roucheng Ou +5

In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural language attack surface unique t…

cs.CL2026

From Training to Generalization: Improving Moral Reasoning Through Pragmatic Inference

Guangliang Liu, Xi Chen, Bocheng Chen +3

Although moral reasoning has emerged as a promising research direction for large language models (LLMs), a persistent generalization challenge remains: LLMs often achieve strong pe…

cs.CL2026

Can Large Language Models Handle Discourse Particles? A Case Study of Colloquial Malay

Mariah Al Giptiah Binte Yusoff, Jakin Tan, Bocheng Chen +2

Discourse particles, such as \textit{well} and \textit{kind of}, are crucial components that enable LLMs to ``speak'' more like humans. They are used to convey emotions, intentions…

cs.CL2026

Learning to Diagnose and Correct Moral Errors: Beyond Shallow Heuristics in Moral Alignment

Bocheng Chen, Xi Chen, Han Zi +5

Existing approaches to moral value alignment are primarily set out to align LLMs' generation with the distributions of morally appropriate language, which has seen good progress. H…

cs.AI2026

Deactivating Refusal Triggers: Understanding and Mitigating Overrefusal in Safety Alignment

Zhiyu Xue, Zimo Qi, Guangliang Liu +2

Safety alignment aims to ensure that large language models (LLMs) refuse harmful requests by post-training on harmful queries paired with refusal answers. Although safety alignment…

cs.CL2025

Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes

Guangliang Liu, Bocheng Chen, Han Zi +2

Moral alignment has emerged as a widely adopted approach for regulating the behavior of pretrained language models (PLMs), typically through fine-tuning on curated datasets. Gender…