3 papers
cs.CL2026
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
Han Jiang, Dongyao Zhu, Xiaoyuan Yi +3
In-Context Learning has shown great potential for aligning Large Language Models (LLMs) with human values, helping reduce harmful outputs and accommodate diverse preferences withou…
cs.CL2025
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
Han Jiang, Xiaoyuan Yi, Zhihua Wei +3
Warning: Contains harmful model outputs. Despite significant advancements, the propensity of Large Language Models (LLMs) to generate harmful and unethical content poses critical c…
cs.CL2025
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
Yu Li, Han Jiang, Zhihua Wei
With the widespread adoption of Large Language Models (LLMs), jailbreak attacks have become an increasingly pressing safety concern. While safety-aligned LLMs can effectively defen…