artificial intelligence

PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments

arXiv:2607.27354

summary

PAUSE is a benchmark that evaluates personal AI assistants on their ability to manage persistent user state, respect configurations and permissions, and coordinate actions across multiple services in realistic, multi‑turn interactions.

Abstract

Personal AI assistants are increasingly deployed as task-oriented, tool-augmented agents that operate within unified service environments to support everyday user activities. In realistic settings, such assistants must reason over persistent user state, respect user-specific configurations and permissions, and sustain long-horizon, constraint-aware interactions across multiple services. Existing benchmarks, however, often fragment service contexts or abstract away user state, limiting their ability to evaluate user-centric personal assistant behavior in realistic service settings. We introduce PAUSE, a user-centric benchmark for evaluating personal AI assistants in stateful, service-integrated environments. PAUSE captures core challenges of real-world assistant deployment by requiring agents to coordinate actions across heterogeneous user-owned resources while maintaining consistency with environment state, authorization constraints over multi-turn interactions. The benchmark incorporates explicit user-agent interaction via realistic user simulation, enabling evaluation beyond static tool execution. To support principled and reproducible evaluation, PAUSE adopts a multi-regime evaluation framework aligned with task characteristics. Open-ended service management tasks are assessed using semantic and trajectory-level behavioral metrics, while constraint-intensive tasks admit deterministic, state-based verification. Benchmark results show that even state-of-the-art proprietary models fail to reach 70% task completion on scenarios requiring stateful reasoning and configuration awareness, revealing consistent and interpretable failure patterns. Finally, we present a user-centric synthesis pipeline that enables scalable generation of coherent service environments, user configurations, and reliably annotated tasks, supporting benchmark extensibility and future research.

Topics & keywords

#personal assistants#benchmarking#stateful reasoning#service integration#user simulationtask-oriented dialoguetool-augmented agentsmulti-turn interactionauthorization constraintsenvironment state
PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments · wovepaper