3 papers
cs.CL2026
Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities
Xiangxu Zhang, Jiamin Wang, Qinlin Zhao +6
As LLMs become increasingly integrated into human society, evaluating their orientations on human values from social science has drawn growing attention. Nevertheless, it is still…
cs.AI2025
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
Hanze Guo, Jing Yao, Xiao Zhou +2
As large language models (LLMs) become increasingly integrated into applications serving users across diverse cultures, communities and demographics, it is critical to align LLMs w…
cs.AI2025
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
Yanxu Zhu, Shitong Duan, Xiangxu Zhang +7
Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despit…