4 papers · 1 filter
Identifying the Unknown: Prompt-Free Open Vocabulary Anomaly Recognition for Robot-Object Interaction
Philipp Allgeuer, Jan-Gerrit Habekost, Stefan Wermter
Robots operating in real-world environments must in general be able to recognize previously unseen objects. As robotic systems move toward open-world autonomy, there is a growing,…
Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders
Sergio Lanza, Jae Hee Lee, Stefan Wermter
Vision Language Models (VLMs) have demonstrated impressive performance in tasks requiring joint understanding of images and text, such as image captioning and Visual Question Answe…
Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation
Hassan Ali, Doreen Jirak, Luca Müller +1
Gesture recognition research, unlike NLP, continues to face acute data scarcity, with progress constrained by the need for costly human recordings or image processing approaches th…
FabuLight-ASD: Unveiling Speech Activity via Body Language
Hugo Carneiro, Stefan Wermter
Active speaker detection (ASD) in multimodal environments is crucial for various applications, from video conferencing to human-robot interaction. This paper introduces FabuLight-A…