3 papers
cs.CV2026
Occlusion-Robust Multimodal Emotion Recognition in VR via Fusion of Facial Images and EMG
Birgit Nierula, Karam Tomotaki-Dawoud, Mert Akguel +5
Head-mounted displays (HMDs) fundamentally limit emotion recognition in virtual reality (VR): by occluding the upper face, they render conventional image-based facial expression an…
cs.CV2026
Event-Level Detection of Surgical Instrument Handovers in Videos with Interpretable Vision Models
Katerina Katsarou, George Zountsas, Karam Tomotaki-Dawoud +6
Reliable monitoring of surgical instrument exchanges is essential for maintaining procedural efficiency and patient safety in the operating room. Automatic detection of instrument…
cs.CV2026
Monocular 3D Object Position Estimation with VLMs for Human-Robot Interaction
Ari Wahl, Dorian Gawlinski, David Przewozny +3
Pre-trained general-purpose Vision-Language Models (VLM) hold the potential to enhance intuitive human-machine interactions due to their rich world knowledge and 2D object detectio…