4 papers
Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models
Bocheng Chen, Han Zi, Roucheng Ou +5
In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural language attack surface unique t…
Learning to Diagnose and Correct Moral Errors: Beyond Shallow Heuristics in Moral Alignment
Bocheng Chen, Xi Chen, Han Zi +5
Existing approaches to moral value alignment are primarily set out to align LLMs' generation with the distributions of morally appropriate language, which has seen good progress. H…
From Training to Generalization: Improving Moral Reasoning Through Pragmatic Inference
Guangliang Liu, Xi Chen, Bocheng Chen +3
Although moral reasoning has emerged as a promising research direction for large language models (LLMs), a persistent generalization challenge remains: LLMs often achieve strong pe…
Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
Guangliang Liu, Bocheng Chen, Han Zi +2
Moral alignment has emerged as a widely adopted approach for regulating the behavior of pretrained language models (PLMs), typically through fine-tuning on curated datasets. Gender…