5 papers
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
Esma Balkır, Alice Pernthaller, Marco Basaldella +2
Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluation increasingly relies on generation tas…
Value of Information: A Framework for Human-Agent Communication
Yijiang River Dong, Tiancheng Hu, Zheng Hui +4
Large Language Model (LLM) agents deployed for real-world tasks face a fundamental dilemma: user requests are underspecified, yet agents must decide whether to act on incomplete in…
Steer Model beyond Assistant: Controlling System Prompt Strength via Contrastive Decoding
Yijiang River Dong, Tiancheng Hu, Zheng Hui +1
Large language models excel at complex instructions yet struggle to deviate from their helpful assistant persona, as post-training instills strong priors that resist conflicting in…
iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to News
Tiancheng Hu, Nigel Collier
Understanding how individuals perceive and react to information is fundamental for advancing social and behavioral sciences and developing human-centered AI systems. Current approa…
When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning
Yijiang River Dong, Tiancheng Hu, Yinhong Liu +2
While Reinforcement Learning from Human Feedback (RLHF) is widely used to align Large Language Models (LLMs) with human preferences, it typically assumes homogeneous preferences ac…