works on

From the 1 of 38 linked papers with an AI index.

activity
20242026
most citedParaphrase Types Elicit Prompt Engineering Capabilities

5 citations · 5 across the 11 of their papers we have counts for

collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

Kia-Jüng Yang, Dominik Meier, Jiachen Zhao +2

Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may complicate control mechanisms suc…

cs.AI2026

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling

Florian Valentin Wunderlich, Lars Benedikt Kaesberg, Jan Philip Wahle +2

Advances in inference methods have enabled language models to improve their predictions without additional training. These methods often prioritize raw performance over cost-effect…

cs.AI2026

Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym

Lars Benedikt Kaesberg, Tianyu Yang, Niklas Bauer +3

Spatial reasoning is central to navigation and robotics, yet measuring model capabilities on these tasks remains difficult. Existing benchmarks evaluate models in a one-shot settin…

cs.AI2025

ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents

Tianyu Yang, Terry Ruas, Yijun Tian +3

Vision-language models (VLMs) excel at interpreting text-rich images but struggle with long, visually complex documents that demand analysis and integration of information spread a…

cs.AI2025

SPaRC: A Spatial Pathfinding Reasoning Challenge

Lars Benedikt Kaesberg, Jan Philip Wahle, Terry Ruas +1

Existing reasoning datasets saturate and fail to test abstract, multi-step problems, especially pathfinding and complex rule constraint satisfaction. We introduce SPaRC (Spatial Pa…

cs.AI2025

You need to MIMIC to get FAME: Solving Meeting Transcript Scarcity with a Multi-Agent Conversations

Frederic Kirstein, Muneeb Khan, Jan Philip Wahle +2

Meeting summarization suffers from limited high-quality data, mainly due to privacy restrictions and expensive collection processes. We address this gap with FAME, a dataset of 500…