works on

From the 1 of 40 linked papers with an AI index.

activity
20242026
collaborators

36 papers

cs.CL2026

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa +3

The paper introduces ParliamentBench, an open-source benchmark based on the Secret Hitler game, to evaluate large language model agents on deception, persuasion, and reasoning unde…

cs.MA2026

Misinformation Propagation in Benign Multi-Agent Systems

Jonas Becker, Jan Philip Wahle, Terry Ruas +1

Multi-agent systems, in which multiple large language model agents solve problems through turn-based interaction, are increasingly deployed in high-stakes settings such as medical…

cs.AI2026

Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

Kia-Jüng Yang, Dominik Meier, Jiachen Zhao +2

Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may complicate control mechanisms suc…

cs.IR2026

Aspect-Aware Content-Based Recommendations for Mathematical Research Papers

Ankit Satpute, André Greiner-Petter, Noah Gießing +4

Content-based research paper recommendation (CbRPR) has seen advances in computer science and biomedicine, but remains unexplored for mathematics, where paper relatedness is more c…

cs.AI2026

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling

Florian Valentin Wunderlich, Lars Benedikt Kaesberg, Jan Philip Wahle +2

Advances in inference methods have enabled language models to improve their predictions without additional training. These methods often prioritize raw performance over cost-effect…

cs.CL2026

Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains

Finn Schmidt, Jan Philip Wahle, Terry Ruas +1

Automatic evaluation metrics are central to the development of machine translation systems, yet their robustness under domain shift remains unclear. Most metrics are developed on t…