1.8k citations · 1.8k across the 5 of their papers we have counts for
3 papers · 1 filter
Transcription-Enriched Joint Embeddings for Spoken Descriptions of Images and Videos
Benet Oriol, Jordi Luque, Ferran Diego +1
In this work, we propose an effective approach for training unique embedding representations by combining three simultaneous modalities: image and spoken and textual narratives. Th…
Seeing and Hearing Egocentric Actions: How Much Can We Learn?
Alejandro Cartas, Jordi Luque, Petia Radeva +2
Our interaction with the world is an inherently multimodal experience. However, the understanding of human-to-object interactions has historically been addressed focusing on a sing…
How Much Does Audio Matter to Recognize Egocentric Object Interactions?
Alejandro Cartas, Jordi Luque, Petia Radeva +2
Sounds are an important source of information on our daily interactions with objects. For instance, a significant amount of people can discern the temperature of water that it is b…