most citedThe Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment

1 citations · 1 across the 1 of their papers we have counts for

collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization

HyunJin Kim, Xiaoyuan Yi, Jing Yao +4

The recent leap in AI capabilities, driven by big generative models, has sparked the possibility of achieving Artificial General Intelligence (AGI) and further triggered discussion…

cs.AI2026

BotzoneBench: Scalable LLM Evaluation via Graded AI Anchors

Lingfeng Li, Yunlong Lu, Yuefei Zhang +7

Large Language Models (LLMs) are increasingly deployed in interactive environments requiring strategic decision-making, yet systematic evaluation of these capabilities remains chal…

cs.AI2026

MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions

Yanxu Zhu, Shitong Duan, Xiangxu Zhang +7

Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despit…

cs.AI2025

Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models

Hanze Guo, Jing Yao, Xiao Zhou +2

As large language models (LLMs) become increasingly integrated into applications serving users across diverse cultures, communities and demographics, it is critical to align LLMs w…

cs.AI2025

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values

Jing Yao, Xiaoyuan Yi, Shitong Duan +8

As Large Language Models (LLMs) achieve remarkable breakthroughs, aligning their values with humans has become imperative for their responsible development and customized applicati…

cs.AI2024

On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models

Xinpeng Wang, Shitong Duan, Xiaoyuan Yi +7

Big models have achieved revolutionary breakthroughs in the field of AI, but they might also pose potential concerns. Addressing such concerns, alignment technologies were introduc…