From the 1 of 27 linked papers with an AI index.
27 papers
Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa +3
The paper introduces ParliamentBench, an open-source benchmark based on the Secret Hitler game, to evaluate large language model agents on deception, persuasion, and reasoning unde…
Misinformation Propagation in Benign Multi-Agent Systems
Jonas Becker, Jan Philip Wahle, Terry Ruas +1
Multi-agent systems, in which multiple large language model agents solve problems through turn-based interaction, are increasingly deployed in high-stakes settings such as medical…
Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling
Florian Valentin Wunderlich, Lars Benedikt Kaesberg, Jan Philip Wahle +2
Advances in inference methods have enabled language models to improve their predictions without additional training. These methods often prioritize raw performance over cost-effect…
DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis
Lung-Hao Lee, Liang-Chih Yu, Natalia Loukashevich +13
Aspect-Based Sentiment Analysis (ABSA) focuses on extracting sentiment at a fine-grained aspect level and has been widely applied across real-world domains. However, existing ABSA…
Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains
Finn Schmidt, Jan Philip Wahle, Terry Ruas +1
Automatic evaluation metrics are central to the development of machine translation systems, yet their robustness under domain shift remains unclear. Most metrics are developed on t…
Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym
Lars Benedikt Kaesberg, Tianyu Yang, Niklas Bauer +3
Spatial reasoning is central to navigation and robotics, yet measuring model capabilities on these tasks remains difficult. Existing benchmarks evaluate models in a one-shot settin…