41 citations · 123 across the 9 of their papers we have counts for
13 papers
Cyclic Co-Learning of Sounding Object Visual Grounding and Sound Separation
Yapeng Tian, Di Hu, Chenliang Xu
There are rich synchronized audio and visual events in our daily life. Inside the events, audio scenes are associated with the corresponding visual objects; meanwhile, sounding obj…
Temporal Relational Modeling with Self-Supervision for Action Segmentation
Dong Wang, Di Hu, Xingjian Li +1
Temporal relational modeling in video is essential for human action understanding, such as action recognition and action segmentation. Although Graph Convolution Networks (GCNs) ha…
Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching
Di Hu, Rui Qian, Minyue Jiang +5
Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a…
Multiple Sound Sources Localization from Coarse to Fine
Rui Qian, Di Hu, Heinrich Dinkel +3
How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this proble…
Ambient Sound Helps: Audiovisual Crowd Counting in Extreme Conditions
Di Hu, Lichao Mou, Qingzhong Wang +4
Visual crowd counting has been recently studied as a way to enable people counting in crowd scenes from images. Albeit successful, vision-based crowd counting approaches could fail…
Cross-Task Transfer for Geotagged Audiovisual Aerial Scene Recognition
Di Hu, Xuhong Li, Lichao Mou +5
Aerial scene recognition is a fundamental task in remote sensing and has recently received increased interest. While the visual information from overhead images with powerful model…