1 paper · 1 filter
Mark Hamilton, Andrew Zisserman, John R. Hershey +1
We present DenseAV, a novel dual encoder grounding architecture that learns high-resolution, semantically meaningful, and audio-visually aligned features solely through watching vi…