Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
cs.AI2026
Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition
Muhammad Ahmed Mohsin, Ahsan Bilal, Muhammad Umer +1
Large language models exhibit sycophancy, the tendency to shift their stated positions toward perceived user preferences or authority cues regardless of evidence. Standard alignmen…
cs.AI2026
How Well Do Multimodal Models Reason on ECG Signals?
Maxwell A. Xu, Harish Haresamudram, Catherine W. Liu +11
While multimodal large language models offer a promising solution to the "black box" nature of health AI by generating interpretable reasoning traces, verifying the validity of the…