1 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 1 cited
Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language
Mark Hamilton, Andrew Zisserman, John R. Hershey +1
We present DenseAV, a novel dual encoder grounding architecture that learns high-resolution, semantically meaningful, and audio-visually aligned features solely through watching vi…
cs.SD2023★ 1 cited
Large-Scale Automatic Audiobook Creation
Brendan Walsh, Mark Hamilton, Greg Newby +8
An audiobook can dramatically improve a work of literature's accessibility and improve reader engagement. However, audiobooks can take hundreds of hours of human effort to create,…