activity
20242026
collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning

Zhaowei Zhang, Xiaohan Liu, Xuekai Zhu +6

Reinforcement learning with verifiable rewards (RLVR) has achieved remarkable success in logical reasoning tasks, yet whether large language model (LLM) alignment requires fundamen…

cs.AI2025

Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models

Hanze Guo, Jing Yao, Xiao Zhou +2

As large language models (LLMs) become increasingly integrated into applications serving users across diverse cultures, communities and demographics, it is critical to align LLMs w…

cs.AI2025

MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions

Yanxu Zhu, Shitong Duan, Xiangxu Zhang +7

Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despit…

cs.AI2025

Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization

HyunJin Kim, Xiaoyuan Yi, Jing Yao +4

The recent leap in AI capabilities, driven by big generative models, has sparked the possibility of achieving Artificial General Intelligence (AGI) and further triggered discussion…

cs.AI2025

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values

Jing Yao, Xiaoyuan Yi, Shitong Duan +8

As Large Language Models (LLMs) achieve remarkable breakthroughs, aligning their values with humans has become imperative for their responsible development and customized applicati…