1 paper
Wenhuan Lu, Xinyue Song, Wenjun Ke +3
Audio grounding, or speech-driven open-set object detection, aims to localize and identify objects directly from speech, enabling generalization beyond predefined categories. This…