collaborators

8 papers

cs.CL2023

AraMUS: Pushing the Limits of Data and Model Scale for Arabic Natural Language Processing

Asaad Alghamdi, Xinyu Duan, Wei Jiang +9

Developing monolingual large Pre-trained Language Models (PLMs) is shown to be very successful in handling different tasks in Natural Language Processing (NLP). In this work, we pr…

cs.CL2023

Measuring the Knowledge Acquisition-Utilization Gap in Pretrained Language Models

Amirhossein Kazemnejad, Mehdi Rezagholizadeh, Prasanna Parthasarathi +1

While pre-trained language models (PLMs) have shown evidence of acquiring vast amounts of knowledge, it remains unclear how much of this parametric knowledge is actually usable in…

eess.AS20231 cited

On the Transferability of Whisper-based Representations for "In-the-Wild" Cross-Task Downstream Speech Applications

Vamsikrishna Chemudupati, Marzieh Tahaei, Heitor Guimaraes +5

Large self-supervised pre-trained speech models have achieved remarkable success across various speech-processing tasks. The self-supervised training of these models leads to unive…

eess.AS2023

An Exploration into the Performance of Unsupervised Cross-Task Speech Representations for "In the Wild'' Edge Applications

Heitor Guimarães, Arthur Pimentel, Anderson Avila +2

Unsupervised speech models are becoming ubiquitous in the speech and machine learning communities. Upstream models are responsible for learning meaningful representations from raw…

cs.IR20233 cited

Simple Yet Effective Neural Ranking and Reranking Baselines for Cross-Lingual Information Retrieval

Jimmy Lin, David Alfonso-Hermelo, Vitor Jeronymo +8

The advent of multilingual language models has generated a resurgence of interest in cross-lingual information retrieval (CLIR), which is the task of searching documents in one lan…

eess.AS2023

RobustDistiller: Compressing Universal Speech Representations for Enhanced Environment Robustness

Heitor R. Guimarães, Arthur Pimentel, Anderson R. Avila +3

Self-supervised speech pre-training enables deep neural network models to capture meaningful and disentangled factors from raw waveform signals. The learned universal speech repres…