1 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 1 cited
Adaptive Length Image Tokenization via Recurrent Allocation
Shivam Duggal, Phillip Isola, Antonio Torralba +1
Current vision systems typically assign fixed-length representations to images, regardless of the information content. This contrasts with human intelligence - and even large langu…
cs.CV2024★ 1 cited
Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language
Mark Hamilton, Andrew Zisserman, John R. Hershey +1
We present DenseAV, a novel dual encoder grounding architecture that learns high-resolution, semantically meaningful, and audio-visually aligned features solely through watching vi…