most citedOn the Transferability of Whisper-based Representations for "In-the-Wild" Cross-Task Downstream Speech Applications

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

eess.AS2024

An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning

Heitor R. Guimarães, Arthur Pimentel, Anderson R. Avila +3

Self-supervised speech representation learning enables the extraction of meaningful features from raw waveforms. These features can then be efficiently used across multiple downstr…

eess.AS2023

On the Impact of Quantization and Pruning of Self-Supervised Speech Models for Downstream Speech Recognition Tasks "In-the-Wild''

Arthur Pimentel, Heitor Guimarães, Anderson R. Avila +2

Recent advances with self-supervised learning have allowed speech recognition systems to achieve state-of-the-art (SOTA) word error rates (WER) while requiring only a fraction of t…

eess.AS2023

VIC-KD: Variance-Invariance-Covariance Knowledge Distillation to Make Keyword Spotting More Robust Against Adversarial Attacks

Heitor R. Guimarães, Arthur Pimentel, Anderson Avila +1

Keyword spotting (KWS) refers to the task of identifying a set of predefined words in audio streams. With the advances seen recently with deep neural networks, it has become a popu…

eess.AS20231 cited

On the Transferability of Whisper-based Representations for "In-the-Wild" Cross-Task Downstream Speech Applications

Vamsikrishna Chemudupati, Marzieh Tahaei, Heitor Guimaraes +5

Large self-supervised pre-trained speech models have achieved remarkable success across various speech-processing tasks. The self-supervised training of these models leads to unive…

eess.AS2023

An Exploration into the Performance of Unsupervised Cross-Task Speech Representations for "In the Wild'' Edge Applications

Heitor Guimarães, Arthur Pimentel, Anderson Avila +2

Unsupervised speech models are becoming ubiquitous in the speech and machine learning communities. Upstream models are responsible for learning meaningful representations from raw…

eess.AS2023

RobustDistiller: Compressing Universal Speech Representations for Enhanced Environment Robustness

Heitor R. Guimarães, Arthur Pimentel, Anderson R. Avila +3

Self-supervised speech pre-training enables deep neural network models to capture meaningful and disentangled factors from raw waveform signals. The learned universal speech repres…