20 papers
Towards High-Level Semantic Intelligence
Xiujie Song, Gefei Yang, Yining You +6
Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development of AI reveals a clear traje…
LatentRevise: Learning from Zero-Hit Reasoning
Yiqiu Guo, Xueting Han, Qi Jia +2
Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by hard prompts on which correct trajectories have low probability, so sampling misses them within a practical…
Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating
Sicheng Wang, Xiangyang Zhu, Han Wang +6
Prior work has shown that fine-tuning large language models on malicious or incorrect outputs in narrow domains can induce broad misalignment and harmful behavior, a phenomenon kno…
UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities
Qi Jia, Haodong Zhao, Dun Pei +7
Benchmarking large language models (LLMs) and agents in multi-turn interactive scenarios is essential for understanding their practical capabilities. However, existing evaluation p…
SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond
Xiangyang Zhu, Yuan Tian, Qi Jia +14
The success of large language models (LLMs) in scientific domains has heightened safety concerns, prompting numerous benchmarks to evaluate their scientific safety. Existing benchm…
VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models
Yunhao Li, Sijing Wu, Zhilin Gao +5
Large multimodal models (LMMs) have demonstrated outstanding capabilities in various visual perception tasks, which has in turn made the evaluation of LMMs significant. However, th…