2 citations · 2 across the 1 of their papers we have counts for
5 papers
Efficient Keyword Spotting by capturing long-range interactions with Temporal Lambda Networks
Biel Tura, Santiago Escuder, Ferran Diego +2
Models based on attention mechanisms have shown unprecedented speech recognition performance. However, they are computationally expensive and unnecessarily complex for keyword spot…
Enabling Zero-shot Multilingual Spoken Language Translation with Language-Specific Encoders and Decoders
Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa +1
Current end-to-end approaches to Spoken Language Translation (SLT) rely on limited training resources, especially for multilingual settings. On the other hand, Multilingual Neural…
Seeing and Hearing Egocentric Actions: How Much Can We Learn?
Alejandro Cartas, Jordi Luque, Petia Radeva +2
Our interaction with the world is an inherently multimodal experience. However, the understanding of human-to-object interactions has historically been addressed focusing on a sing…
How Much Does Audio Matter to Recognize Egocentric Object Interactions?
Alejandro Cartas, Jordi Luque, Petia Radeva +2
Sounds are an important source of information on our daily interactions with objects. For instance, a significant amount of people can discern the temperature of water that it is b…
Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion
Joan Serrà, Santiago Pascual, Carlos Segura
End-to-end models for raw audio generation are a challenge, specially if they have to work with non-parallel data, which is a desirable setup in many situations. Voice conversion,…