1 paper
Wenhao Yang, Jianguo Wei, Wenhuan Lu +1
Grounding objects in images using visual cues is a well-established approach in computer vision, yet the potential of audio as a modality for object recognition and grounding remai…