41 citations · 123 across the 9 of their papers we have counts for
14 papers · 1 filter
Robust Cross-Modal Knowledge Distillation for Unconstrained Videos
Wenke Xia, Xingjian Li, Andong Deng +3
Cross-modal distillation has been widely used to transfer knowledge across different modalities, enriching the representation of the target unimodal one. Recent studies highly rela…
Cyclic Co-Learning of Sounding Object Visual Grounding and Sound Separation
Yapeng Tian, Di Hu, Chenliang Xu
There are rich synchronized audio and visual events in our daily life. Inside the events, audio scenes are associated with the corresponding visual objects; meanwhile, sounding obj…
Temporal Relational Modeling with Self-Supervision for Action Segmentation
Dong Wang, Di Hu, Xingjian Li +1
Temporal relational modeling in video is essential for human action understanding, such as action recognition and action segmentation. Although Graph Convolution Networks (GCNs) ha…
Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching
Di Hu, Rui Qian, Minyue Jiang +5
Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a…
Multiple Sound Sources Localization from Coarse to Fine
Rui Qian, Di Hu, Heinrich Dinkel +3
How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this proble…
Ambient Sound Helps: Audiovisual Crowd Counting in Extreme Conditions
Di Hu, Lichao Mou, Qingzhong Wang +4
Visual crowd counting has been recently studied as a way to enable people counting in crowd scenes from images. Albeit successful, vision-based crowd counting approaches could fail…