activity
20242026
most citedAI Evaluation Should Require Standardized Item-Level Data Releases

1 citations · 2 across the 12 of their papers we have counts for

collaborators
Showing cs.AIShow all

9 papers · 1 filter

cs.AI20261 cited

AI Evaluation Should Require Standardized Item-Level Data Releases

Han Jiang, Susu Zhang, Dongyao Zhu +6

This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified it…

cs.AI2026

Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning

Zhaowei Zhang, Xiaohan Liu, Xuekai Zhu +6

Reinforcement learning with verifiable rewards (RLVR) has achieved remarkable success in logical reasoning tasks, yet whether large language model (LLM) alignment requires fundamen…

cs.AI2026

Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization

HyunJin Kim, Xiaoyuan Yi, Jing Yao +4

The recent leap in AI capabilities, driven by big generative models, has sparked the possibility of achieving Artificial General Intelligence (AGI) and further triggered discussion…

cs.AI2026

MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions

Yanxu Zhu, Shitong Duan, Xiangxu Zhang +7

Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despit…

cs.AI2025

On the Dynamics of Multi-Agent LLM Communities Driven by Value Diversity

Muhua Huang, Qinlin Zhao, Xiaoyuan Yi +1

As Large Language Models (LLM) based multi-agent systems become increasingly prevalent, the collective behaviors, e.g., collective intelligence, of such artificial communities have…

cs.AI2025

Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models

Hanze Guo, Jing Yao, Xiao Zhou +2

As large language models (LLMs) become increasingly integrated into applications serving users across diverse cultures, communities and demographics, it is critical to align LLMs w…