4 papers
Intent Laundering: AI Safety Datasets Are Not What They Seem
Shahriar Golchin, Marc Wetter
We systematically evaluate the quality of widely used adversarial safety datasets from two perspectives: in isolation and in practice. In isolation, we examine how well these datas…
EchoChain: A Full-Duplex Benchmark for State-Update Reasoning Under Interruptions
Smit Nautambhai Modi, Gandharv Mahajan, Marc Wetter +1
Real-time voice assistants must revise task state when users interrupt mid-response, but existing spoken-dialog benchmarks largely evaluate turn-based interaction and miss this fai…
Implicit Intelligence -- Evaluating Agents on What Users Don't Say
Ved Sirdeshmukh, Marc Wetter
Real-world requests to AI agents are fundamentally underspecified. Natural human communication relies on shared context and unstated constraints that speakers expect listeners to i…
R-ConstraintBench: Evaluating LLMs on NP-Complete Scheduling
Raj Jain, Marc Wetter
Effective scheduling under tight resource, timing, and operational constraints underpins large-scale planning across sectors such as capital projects, manufacturing, logistics, and…