6 citations · 7 across the 8 of their papers we have counts for
14 papers
A neural prosody encoder for end-ro-end dialogue act classification
Kai Wei, Dillon Knox, Martin Radfar +6
Dialogue act classification (DAC) is a critical task for spoken language understanding in dialogue systems. Prosodic features such as energy and pitch have been shown to be useful…
Context-Aware Transformer Transducer for Speech Recognition
Feng-Ju Chang, Jing Liu, Martin Radfar +4
End-to-end (E2E) automatic speech recognition (ASR) systems often have difficulty recognizing uncommon words, that appear infrequently in the training data. One promising method, t…
Multi-Channel Transformer Transducer for Speech Recognition
Feng-Ju Chang, Martin Radfar, Athanasios Mouchtaris +1
Multi-channel inputs offer several advantages over single-channel, to improve the robustness of on-device speech recognition systems. Recent work on multi-channel transformer, has…
Sample Drop Detection for Distant-speech Recognition with Asynchronous Devices Distributed in Space
Tina Raissi, Santiago Pascual, Maurizio Omologo
In many applications of multi-microphone multi-device processing, the synchronization among different input channels can be affected by the lack of a common clock and isolated drop…
DiPCo -- Dinner Party Corpus
Maarten Van Segbroeck, Ahmed Zaid, Ksenia Kutsenko +7
We present a speech data corpus that simulates a "dinner party" scenario taking place in an everyday home environment. The corpus was created by recording multiple groups of four A…
LOCATA challenge: speaker localization with a planar array
Xinyuan Qian, Andrea Cavallaro, Alessio Brutti +1
This document describes our submission to the 2018 LOCalization And TrAcking (LOCATA) challenge (Tasks 1, 3, 5). We estimate the 3D position of a speaker using the Global Coherence…