activity
20242026
most citedSemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion Detection

1 citations · 1 across the 13 of their papers we have counts for

collaborators

27 papers

cs.CV2026

Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds

Lars Benedikt Kaesberg, Tianyu Yang, Florian Valentin Wunderlich +4

Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and dow…

cs.CL2026

BabelSteering: Multilingual Safety Alignment via English Steering Vectors

Emma V. Stein, Dominik Meier, Terry Ruas +2

Large language models (LLMs) are deployed globally in high-stakes settings, yet most safety research and alignment efforts remain concentrated on English. Thus, users interacting w…

cs.CL2026

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa +3

As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safe…

cs.MA2026

Misinformation Propagation in Benign Multi-Agent Systems

Jonas Becker, Jan Philip Wahle, Terry Ruas +1

Multi-agent systems, in which multiple large language model agents solve problems through turn-based interaction, are increasingly deployed in high-stakes settings such as medical…

cs.AI2026

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling

Florian Valentin Wunderlich, Lars Benedikt Kaesberg, Jan Philip Wahle +2

Advances in inference methods have enabled language models to improve their predictions without additional training. These methods often prioritize raw performance over cost-effect…

cs.CL2026

Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains

Finn Schmidt, Jan Philip Wahle, Terry Ruas +1

Automatic evaluation metrics are central to the development of machine translation systems, yet their robustness under domain shift remains unclear. Most metrics are developed on t…