Sensible Agent: A Framework for Unobtrusive Interaction with Proactive AR Agents
arXiv:2509.09255 · doi:10.1145/3746059.3747748
Abstract
Proactive AR agents promise context-aware assistance, but their interactions often rely on explicit voice prompts or responses, which can be disruptive or socially awkward. We introduce Sensible Agent, a framework designed for unobtrusive interaction with these proactive agents. Sensible Agent dynamically adapts both "what" assistance to offer and, crucially, "how" to deliver it, based on real-time multimodal context sensing. Informed by an expert workshop (n=12) and a data annotation study (n=40), the framework leverages egocentric cameras, multimodal sensing, and Large Multimodal Models (LMMs) to infer context and suggest appropriate actions delivered via minimally intrusive interaction modes. We demonstrate our prototype on an XR headset through a user study (n=10) in both AR and VR scenarios. Results indicate that Sensible Agent significantly reduces perceived interaction effort compared to voice-prompted baseline, while maintaining high usability and achieving higher preference.
References in corpus (26)
- TensorFlow.js: Machine Learning for the Web and Beyond
- Exploring Interactions Between Trust, Anthropomorphism, and Relationship Development in Voice Assistants
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented Reality
- Towards Question-based Recommender Systems
- XAIR: A Framework of Explainable AI in Augmented Reality
- Augmented Object Intelligence with XR-Objects
- Learning to Ask Appropriate Questions in Conversational Recommendation
- Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
- Proactive Human-Machine Conversation with Explicit Conversation Goals
- AiGet: Transforming Everyday Moments into Hidden Knowledge Discovery with AI Assistance on Smart Glasses
- CookAR: Affordance Augmentations in Wearable AR to Support Kitchen Tool Interactions for People with Low Vision
- Human I/O: Towards a Unified Approach to Detecting Situational Impairments
- WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech Interactions
- Proactive Retrieval-based Chatbots based on Relevant Knowledge and Goals
- Towards Conversational Recommendation over Multi-Type Dialogs
- ARLang: An Outdoor Augmented Reality Application for Portuguese Vocabulary Learning
- Asking Clarifying Questions Based on Negative Feedback in Conversational Search
- Balancing Reinforcement Learning Training Experiences in Interactive Information Retrieval
- Agreement-on-the-Line: Predicting the Performance of Neural Networks under Distribution Shift
- Target-Guided Open-Domain Conversation Planning
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- Can AI Prompt Humans? Multimodal Agents Prompt Players' Game Actions and Show Consequences to Raise Sustainability Awareness
- Working in Extended Reality in the Wild: Worker and Bystander Experiences of XR Virtual Displays in Real-World Settings
- PARSE-Ego4D: Personal Action Recommendation Suggestions for Egocentric Videos
- YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
- Learning to Ask Critical Questions for Assisting Product Search