348 citations · 384 across the 7 of their papers we have counts for
4 papers · 1 filter
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
Dionyssos Kounadis-Bastian, Oliver Schrüfer, Anna Derington +4
Speech Emotion Recognition (SER) needs high computational resources to overcome the challenge of substantial annotator disagreement. Today SER is shifting towards dimensional annot…
Speech-based Age and Gender Prediction with Transformers
Felix Burkhardt, Johannes Wagner, Hagen Wierstorf +2
We report on the curation of several publicly available datasets for age and gender prediction. Furthermore, we present experiments to predict age and gender with models based on a…
Referenceless Performance Evaluation of Audio Source Separation using Deep Neural Networks
Emad M. Grais, Hagen Wierstorf, Dominic Ward +2
Current performance evaluation for audio source separation depends on comparing the processed or separated signals with reference signals. Therefore, common performance evaluation…
Multi-Resolution Fully Convolutional Neural Networks for Monaural Audio Source Separation
Emad M. Grais, Hagen Wierstorf, Dominic Ward +1
In deep neural networks with convolutional layers, each layer typically has fixed-size/single-resolution receptive field (RF). Convolutional layers with a large RF capture global i…