6 papers
Robust Audio Tagging under Class-wise Supervision Unreliability
Yuanbo Hou, Zhaoyi Liu, Tong Ye +4
Weakly labeled datasets such as AudioSet have driven recent progress in audio tagging. However, annotation quality varies across sound classes. Labels may be incomplete, ambiguous,…
Geo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context
Yuanbo Hou, Yanru Wu, Qiaoqiao Ren +3
Environmental sound understanding in computational auditory scene analysis (CASA) is often formulated as an audio-only recognition problem. This formulation leaves a persistent dra…
Angle-Optimized Partial Disentanglement for Multimodal Emotion Recognition in Conversation
Xinyi Che, Wenbo Wang, Yuanbo Hou +3
Multimodal Emotion Recognition in Conversation (MERC) aims to enhance emotion understanding by integrating complementary cues from text, audio, and visual modalities. Existing MERC…
Soundscape Captioning using Sound Affective Quality Network and Large Language Model
Yuanbo Hou, Qiaoqiao Ren, Andrew Mitchell +4
We live in a rich and varied acoustic world, which is experienced by individuals or communities as a soundscape. Computational auditory scene analysis, disentangling acoustic scene…
Touch and Tell: Multimodal Decoding of Human Emotions and Social Gestures for Robots
Qiaoqiao Ren, Remko Proesmans, Yuanbo Hou +2
Human emotions are complex and can be conveyed through nuanced touch gestures. Previous research has primarily focused on how humans recognize emotions through touch or on identify…
Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot Interaction
Yuanbo Hou, Qiaoqiao Ren, Wenwu Wang +1
Emotion recognition and touch gesture decoding are crucial for advancing human-robot interaction (HRI), especially in social environments where emotional cues and tactile perceptio…