7 citations · 22 across the 29 of their papers we have counts for
8 papers · 1 filter
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
Zhaowei Zhang, Xiaohan Liu, Xuekai Zhu +6
Reinforcement learning with verifiable rewards (RLVR) has achieved remarkable success in logical reasoning tasks, yet whether large language model (LLM) alignment requires fundamen…
On the Dynamics of Multi-Agent LLM Communities Driven by Value Diversity
Muhua Huang, Qinlin Zhao, Xiaoyuan Yi +1
As Large Language Models (LLM) based multi-agent systems become increasingly prevalent, the collective behaviors, e.g., collective intelligence, of such artificial communities have…
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
Hanze Guo, Jing Yao, Xiao Zhou +2
As large language models (LLMs) become increasingly integrated into applications serving users across diverse cultures, communities and demographics, it is critical to align LLMs w…
The Morality of Probability: How Implicit Moral Biases in LLMs May Shape the Future of Human-AI Symbiosis
Eoin O'Doherty, Nicole Weinrauch, Andrew Talone +4
Artificial intelligence (AI) is advancing at a pace that raises urgent questions about how to align machine decision-making with human moral values. This working paper investigates…
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
Han Jiang, Pengda Wang, Xiaoyuan Yi +2
Social sciences have accumulated a rich body of theories and methodologies for investigating the human mind and behaviors, while offering valuable insights into the design and unde…
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
Yanxu Zhu, Shitong Duan, Xiangxu Zhang +7
Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despit…