activity
20002021
most citedDynamic Evaluation of Neural Sequence Models

60 citations · 170 across the 14 of their papers we have counts for

collaborators
Showing eess.ASShow all

11 papers · 1 filter

eess.AS202110 cited

Automatic audiovisual synchronisation for ultrasound tongue imaging

Aciel Eshky, Joanne Cleland, Manuel Sam Ribeiro +3

Ultrasound tongue imaging is used to visualise the intra-oral articulators during speech production. It is utilised in a range of applications, including speech and language therap…

eess.AS2021

Silent versus modal multi-speaker speech recognition from ultrasound and video

Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond +1

We investigate multi-speaker speech recognition from ultrasound images of the tongue and video images of the lips. We train our systems on imaging data from modal speech, and evalu…

eess.AS202127 cited

Exploiting ultrasound tongue imaging for the automatic detection of speech articulation errors

Manuel Sam Ribeiro, Joanne Cleland, Aciel Eshky +2

Speech sound disorders are a common communication impairment in childhood. Because speech disorders can negatively affect the lives and the development of children, clinical interv…

eess.AS2021

Train your classifier first: Cascade Neural Networks Training from upper layers to lower layers

Shucong Zhang, Cong-Thanh Do, Rama Doddipatla +3

Although the lower layers of a deep neural network learn features which are transferable across datasets, these layers are not transferable within the same dataset. That is, in gen…

eess.AS2020

TaL: a synchronised multi-speaker corpus of ultrasound tongue imaging, audio, and lip videos

Manuel Sam Ribeiro, Jennifer Sanger, Jing-Xuan Zhang +4

We present the Tongue and Lips corpus (TaL), a multi-speaker corpus of audio, ultrasound tongue imaging, and lip videos. TaL consists of two parts: TaL1 is a set of six recording s…

eess.AS2020

Word Error Rate Estimation Without ASR Output: e-WER2

Ahmed Ali, Steve Renals

Measuring the performance of automatic speech recognition (ASR) systems requires manually transcribed data in order to compute the word error rate (WER), which is often time-consum…