Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System Design
arXiv:2501.13443 · doi:10.1145/3706598.3714161
Abstract
The recent surge in artificial intelligence, particularly in multimodal processing technology, has advanced human-computer interaction, by altering how intelligent systems perceive, understand, and respond to contextual information (i.e., context awareness). Despite such advancements, there is a significant gap in comprehensive reviews examining these advances, especially from a multimodal data perspective, which is crucial for refining system design. This paper addresses a key aspect of this gap by conducting a systematic survey of data modality-driven Vision-based Multimodal Interfaces (VMIs). VMIs are essential for integrating multimodal data, enabling more precise interpretation of user intentions and complex interactions across physical and digital environments. Unlike previous task- or scenario-driven surveys, this study highlights the critical role of the visual modality in processing contextual information and facilitating multimodal interaction. Adopting a design framework moving from the whole to the details and back, it classifies VMIs across dimensions, providing insights for developing effective, context-aware systems.
The ACM CHI Conference on Human Factors in Computing Systems 2025 (CHI 2025)
References in corpus (18)
- SMOTE: Synthetic Minority Over-sampling Technique
- Augmented Reality and Robotics: A Survey and Taxonomy for AR-enhanced Human-Robot Interaction and Robotic Interfaces
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text Data
- Mathemyths: Leveraging Large Language Models to Teach Mathematical Language through Child-AI Co-Creative Storytelling
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented Reality
- GPT-4 Enhanced Multimodal Grounding for Autonomous Driving: Leveraging Cross-Modal Attention with Large Language Models
- Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for Radiology
- VirtuWander: Enhancing Multi-modal Interaction for Virtual Tour Guidance through Large Language Models
- From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data With Large Language Models
- Deep Thermal Imaging: Proximate Material Type Recognition in the Wild through Deep Learning of Spatial Surface Temperature Patterns
- Exploring Large-Scale Language Models to Evaluate EEG-Based Multimodal Data for Mental Health
- Leveraging driver vehicle and environment interaction: Machine learning using driver monitoring cameras to detect drunk driving
- Geo-Context Aware Study of Vision-Based Autonomous Driving Models and Spatial Video Data
- PaperToPlace: Transforming Instruction Documents into Spatialized and Context-Aware Mixed Reality Experiences
- Gesture-aware Interactive Machine Teaching with In-situ Object Annotations
- RealityEffects: Augmenting 3D Volumetric Videos with Object-Centric Annotation and Dynamic Visual Effects
- Towards Enhanced Context Awareness with Vision-based Multimodal Interfaces
- MicroCam: Leveraging Smartphone Microscope Camera for Context-Aware Contact Surface Sensing