5 papers
Glance-Say: Multimodal Human-Robot Collaboration and Intent Recognition via Sticky Glance
Yuzhi Lai, Shenghai Yuan, Peizheng Li +2
Gaze and speech are promising interaction modalities for individuals with motor impairments, yet robust intent recognition in multi-object environments remains challenging due to m…
FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech
Yuzhi Lai, Shenghai Yuan, Peizheng Li +4
ffective Human-Robot Interaction (HRI) is crucial for enhancing accessibility and usability in real-world robotics applications. However, existing solutions often rely on gesture-…
4th Workshop on Maritime Computer Vision (MaCVi): Challenge Overview
Benjamin Kiefer, Jan Lukas Augustin, Jon MuhoviÄ +52
The 4th Workshop on Maritime Computer Vision (MaCVi) is organized as part of CVPR 2026. This edition features five benchmark challenges with emphasis on both predictive accuracy an…
Lightweight Multi-Frame Integration for Robust YOLO Object Detection in Videos
Yitong Quan, Benjamin Kiefer, Martin Messmer +1
Modern image-based object detection models, such as YOLOv7, primarily process individual frames independently, thus ignoring valuable temporal context naturally present in videos.…
Approximate Supervised Object Distance Estimation on Unmanned Surface Vehicles
Benjamin Kiefer, Yitong Quan, Andreas Zell
Unmanned surface vehicles (USVs) and boats are increasingly important in maritime operations, yet their deployment is limited due to costly sensors and complexity. LiDAR, radar, an…