GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
arXiv:2401.14268 · doi:10.1145/3654777.3676356
Abstract
Virtual assistants have the potential to play an important role in helping users achieves different tasks. However, these systems face challenges in their real-world usability, characterized by inefficiency and struggles in grasping user intentions. Leveraging recent advances in Large Language Models (LLMs), we introduce GptVoiceTasker, a virtual assistant poised to enhance user experiences and task efficiency on mobile devices. GptVoiceTasker excels at intelligently deciphering user commands and executing relevant device interactions to streamline task completion. The system continually learns from historical user commands to automate subsequent usages, further enhancing execution efficiency. Our experiments affirm GptVoiceTasker's exceptional command interpretation abilities and the precision of its task automation module. In our user study, GptVoiceTasker boosted task efficiency in real-world scenarios by 34.85%, accompanied by positive participant feedback. We made GptVoiceTasker open-source, inviting further research into LLMs utilization for diverse tasks through prompt engineering and leveraging user usage data to improve efficiency.
This paper has been accepted by UIST 2024
References in corpus (4)
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented Reality
- MadDroid: Characterising and Detecting Devious Ad Content for Android Apps
- LM4HPC: Towards Effective Language Model Application in High-Performance Computing
- Voicify Your UI: Towards Android App Control with Voice Commands
Cited by in corpus (4)
- From Interaction to Impact: Towards Safer AI Agents Through Understanding and Evaluating Mobile UI Operation Impacts
- Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts
- Harnessing Large Language Model for Virtual Reality Exploration Testing: A Case Study
- Caption: Generating Informative Content Labels for Image Buttons Using Next-Screen Context