activity
20192021
most citedHow Much Does Audio Matter to Recognize Egocentric Object Interactions?

2 citations · 2 across the 1 of their papers we have counts for

collaborators

5 papers

eess.AS2021

Efficient Keyword Spotting by capturing long-range interactions with Temporal Lambda Networks

Biel Tura, Santiago Escuder, Ferran Diego +2

Models based on attention mechanisms have shown unprecedented speech recognition performance. However, they are computationally expensive and unnecessarily complex for keyword spot…

cs.CL2020

Enabling Zero-shot Multilingual Spoken Language Translation with Language-Specific Encoders and Decoders

Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa +1

Current end-to-end approaches to Spoken Language Translation (SLT) rely on limited training resources, especially for multilingual settings. On the other hand, Multilingual Neural…

cs.CV2019

Seeing and Hearing Egocentric Actions: How Much Can We Learn?

Alejandro Cartas, Jordi Luque, Petia Radeva +2

Our interaction with the world is an inherently multimodal experience. However, the understanding of human-to-object interactions has historically been addressed focusing on a sing…

cs.CV20192 cited

How Much Does Audio Matter to Recognize Egocentric Object Interactions?

Alejandro Cartas, Jordi Luque, Petia Radeva +2

Sounds are an important source of information on our daily interactions with objects. For instance, a significant amount of people can discern the temperature of water that it is b…

cs.LG2019

Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion

Joan Serrà, Santiago Pascual, Carlos Segura

End-to-end models for raw audio generation are a challenge, specially if they have to work with non-parallel data, which is a desirable setup in many situations. Voice conversion,…