works on

From the 1 of 27 linked papers with an AI index.

collaborators

27 papers

cs.CL2026

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa +3

The paper introduces ParliamentBench, an open-source benchmark based on the Secret Hitler game, to evaluate large language model agents on deception, persuasion, and reasoning unde…

cs.MA2026

Misinformation Propagation in Benign Multi-Agent Systems

Jonas Becker, Jan Philip Wahle, Terry Ruas +1

Multi-agent systems, in which multiple large language model agents solve problems through turn-based interaction, are increasingly deployed in high-stakes settings such as medical…

cs.AI2026

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling

Florian Valentin Wunderlich, Lars Benedikt Kaesberg, Jan Philip Wahle +2

Advances in inference methods have enabled language models to improve their predictions without additional training. These methods often prioritize raw performance over cost-effect…

cs.CL2026

DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis

Lung-Hao Lee, Liang-Chih Yu, Natalia Loukashevich +13

Aspect-Based Sentiment Analysis (ABSA) focuses on extracting sentiment at a fine-grained aspect level and has been widely applied across real-world domains. However, existing ABSA…

cs.CL2026

Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains

Finn Schmidt, Jan Philip Wahle, Terry Ruas +1

Automatic evaluation metrics are central to the development of machine translation systems, yet their robustness under domain shift remains unclear. Most metrics are developed on t…

cs.AI2026

Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym

Lars Benedikt Kaesberg, Tianyu Yang, Niklas Bauer +3

Spatial reasoning is central to navigation and robotics, yet measuring model capabilities on these tasks remains difficult. Existing benchmarks evaluate models in a one-shot settin…