2 papers
cs.CV2025
Learning Visual Affordance from Audio
Lidong Lu, Guo Chen, Zhu Wei +2
We introduce Audio-Visual Affordance Grounding (AV-AG), a new task that segments object interaction regions from action sounds. Unlike existing approaches that rely on textual inst…
eess.AS2025
Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience
Andrew Chang, Chenkai Hu, Ji Qi +5
Group conversations over videoconferencing are a complex social behavior. However, the subjective moments of negative experience, where the conversation loses fluidity or enjoyment…