Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Argument-Based Comparative Question Answering Evaluation Benchmark
Irina Nikishina, Saba Anwar, Nikolay Dolgov +6
In this paper, we aim to solve the problems standing in the way of automatic comparative question answering. To this end, we propose an evaluation framework to assess the quality o…
cs.CL2025
Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
Yilin Geng, Haonan Li, Honglin Mu +5
Large language models (LLMs) are increasingly deployed with hierarchical instruction schemes, where certain instructions (e.g., system-level directives) are expected to take preced…