2 papers
cs.CL2026
Equal Ranking Quality, Different Decisions: Training Order-Consistent LLM Scorers
Markus Frohmann, Mahdiyar Alavi, Elizabeth Lingg +1
Rerankers, reward models and multi-document QA scorers score candidate documents or responses in one LLM prompt, so each score depends on their order. Such scorers are selected on…
cs.IR2026
When Deep Research Agents Stagnate: Enhancing Reasoning with Retrieval-Aware Agent Control
Heydar Soudani, Elizabeth Lingg, Faegheh Hasibi +1
In this paper, we analyze the reasoning trajectories of a variety of DRAs and show that existing agents often suffer from reasoning stagnation: the majority of iterations contribut…