4 citations · 8 across the 20 of their papers we have counts for
4 papers · 2 filters
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
Hanze Guo, Jing Yao, Xiao Zhou +2
As large language models (LLMs) become increasingly integrated into applications serving users across diverse cultures, communities and demographics, it is critical to align LLMs w…
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
Yanxu Zhu, Shitong Duan, Xiangxu Zhang +7
Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despit…
Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization
HyunJin Kim, Xiaoyuan Yi, Jing Yao +4
The recent leap in AI capabilities, driven by big generative models, has sparked the possibility of achieving Artificial General Intelligence (AGI) and further triggered discussion…
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values
Jing Yao, Xiaoyuan Yi, Shitong Duan +8
As Large Language Models (LLMs) achieve remarkable breakthroughs, aligning their values with humans has become imperative for their responsible development and customized applicati…