Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Stay Focused: Problem Drift in Multi-Agent Debate
Jonas Becker, Lars Benedikt Kaesberg, Andreas Stephan +3
Multi-agent debate - multiple instances of large language models discussing problems in turn-based interaction - has shown promise for solving knowledge and reasoning tasks. Howeve…
cs.CL2025
From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks
Andreas Stephan, Dawei Zhu, Matthias AÃenmacher +2
To reduce the need for human annotations, large language models (LLMs) have been proposed as judges of the quality of other candidate models. The performance of LLM judges is typic…
cs.CL2024
Analysing zero-shot temporal relation extraction on clinical notes using temporal consistency
Vasiliki Kougia, Anastasiia Sedova, Andreas Stephan +2
This paper presents the first study for temporal relation extraction in a zero-shot setting focusing on biomedical text. We employ two types of prompts and five LLMs (GPT-3.5, Mixt…