31 citations · 31 across the 6 of their papers we have counts for
6 papers
Exploring compressibility of transformer based text-to-music (TTM) models
Vasileios Moschopoulos, Thanasis Kotsiopoulos, Pablo Peso Parada +7
State-of-the art Text-To-Music (TTM) generative AI models are large and require desktop or server class compute, making them infeasible for deployment on mobile phones. This paper…
Locality enhanced dynamic biasing and sampling strategies for contextual ASR
Md Asif Jalal, Pablo Peso Parada, George Pavlidis +8
Automatic Speech Recognition (ASR) still face challenges when recognizing time-variant rare-phrases. Contextual biasing (CB) modules bias ASR model towards such contextually-releva…
Consistency Based Unsupervised Self-training For ASR Personalisation
Jisi Zhang, Vandana Rajan, Haaris Mehmood +7
On-device Automatic Speech Recognition (ASR) models trained on speech data of a large population might underperform for individuals unseen during training. This is due to a domain…
On-Device Speaker Anonymization of Acoustic Embeddings for ASR based onFlexible Location Gradient Reversal Layer
Md Asif Jalal, Pablo Peso Parada, Jisi Zhang +5
Smart devices serviced by large-scale AI models necessitates user data transfer to the cloud for inference. For speech applications, this means transferring private user informatio…
Towards domain generalisation in ASR with elitist sampling and ensemble knowledge distillation
Rehan Ahmad, Md Asif Jalal, Muhammad Umar Farooq +2
Knowledge distillation has widely been used for model compression and domain adaptation for speech applications. In the presence of multiple teachers, knowledge can easily be trans…
A cross-corpus study on speech emotion recognition
Rosanna Milner, Md Asif Jalal, Raymond W. M. Ng +1
For speech emotion datasets, it has been difficult to acquire large quantities of reliable data and acted emotions may be over the top compared to less expressive emotions displaye…