2 papers
cs.AI2026
Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science Agents
Josefa Lia Stoisser, Marc Boubnovski Martell, Sidsel Boldsen +2
As data-science agents shift from co-pilots to auto-pilots, silent misframing becomes a critical failure mode. Agents quietly commit to plausible but unintended task framings, prod…
cs.AI2026
Measuring Black-Box Confidence via Reasoning Trajectories: Geometry, Coverage, and Verbalization
Marc Boubnovski Martell, Josefa Lia Stoisser, Kaspar Märtens +4
Reliable confidence estimation enables safe deployment of chain-of-thought (CoT) reasoning through text-only APIs. Yet the dominant black-box baseline, self-consistency over K samp…