2 papers
cs.CL2025
Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
Yilin Geng, Haonan Li, Honglin Mu +5
Large language models (LLMs) are increasingly deployed with hierarchical instruction schemes, where certain instructions (e.g., system-level directives) are expected to take preced…
cs.CL2025
Argument-Based Comparative Question Answering Evaluation Benchmark
Irina Nikishina, Saba Anwar, Nikolay Dolgov +6
In this paper, we aim to solve the problems standing in the way of automatic comparative question answering. To this end, we propose an evaluation framework to assess the quality o…