GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented Reality
arXiv:2404.08213 · doi:10.1145/3613904.3642230
Abstract
Voice assistants (VAs) like Siri and Alexa are transforming human-computer interaction; however, they lack awareness of users' spatiotemporal context, resulting in limited performance and unnatural dialogue. We introduce GazePointAR, a fully-functional context-aware VA for wearable augmented reality that leverages eye gaze, pointing gestures, and conversation history to disambiguate speech queries. With GazePointAR, users can ask "what's over there?" or "how do I solve this math problem?" simply by looking and/or pointing. We evaluated GazePointAR in a three-part lab study (N=12): (1) comparing GazePointAR to two commercial systems; (2) examining GazePointAR's pronoun disambiguation across three tasks; (3) and an open-ended phase where participants could suggest and try their own context-sensitive queries. Participants appreciated the naturalness and human-like nature of pronoun-driven queries, although sometimes pronoun use was counter-intuitive. We then iterated on GazePointAR and conducted a first-person diary study examining how GazePointAR performs in-the-wild. We conclude by enumerating limitations and design considerations for future context-aware VAs.
References in corpus (2)
Cited by in corpus (19)
- Augmented Object Intelligence with XR-Objects
- AiGet: Transforming Everyday Moments into Hidden Knowledge Discovery with AI Assistance on Smart Glasses
- Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System Design
- DesignFromX: Empowering Consumer-Driven Design Space Exploration through Feature Composition of Referenced Products
- GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
- Sensible Agent: A Framework for Unobtrusive Interaction with Proactive AR Agents
- MaRginalia: Enabling In-person Lecture Capturing and Note-taking Through Mixed Reality
- Guided Reality: Generating Visually-Enriched AR Task Guidance with LLMs and Vision Models
- Everyday AR through AI-in-the-Loop
- Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes
- PILAR: Personalizing Augmented Reality Interactions with LLM-based Human-Centric and Trustworthy Explanations for Daily Use Cases
- Reality Proxy: Fluid Interactions with Real-World Objects in MR via Abstract Representations
- StepWrite: Adaptive Planning for Speech-Driven Text Generation
- RealitySummary: Exploring On-Demand Mixed Reality Text Summarization and Question Answering using Large Language Models
- Talking Spell: A Wearable System Enabling Real-Time Anthropomorphic Voice Interaction with Everyday Objects
- SpeechLess: Micro-utterance with Personalized Spatial Memory-aware Assistant in Everyday Augmented Reality
- From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality
- Look and Talk: Seamless AI Assistant Interaction with Gaze-Triggered Activation
- Memento: Towards Proactive Visualization of Everyday Memories with Personal Wearable AR Assistant