An evaluation of intrusive instrumental intelligibility metrics
arXiv:1708.06027 · doi:10.1109/TASLP.2018.2856374
Abstract
Instrumental intelligibility metrics are commonly used as an alternative to listening tests. This paper evaluates 12 monaural intrusive intelligibility metrics: SII, HEGP, CSII, HASPI, NCM, QSTI, STOI, ESTOI, MIKNN, SIMI, SIIB, and . In addition, this paper investigates the ability of intelligibility metrics to generalize to new types of distortions and analyzes why the top performing metrics have high performance. The intelligibility data were obtained from 11 listening tests described in the literature. The stimuli included Dutch, Danish, and English speech that was distorted by additive noise, reverberation, competing talkers, pre-processing enhancement, and post-processing enhancement. SIIB and HASPI had the highest performance achieving a correlation with listening test scores on average of and , respectively. The high performance of SIIB may, in part, be the result of SIIBs developers having access to all the intelligibility data considered in the evaluation. The results show that intelligibility metrics tend to perform poorly on data sets that were not used during their development. By modifying the original implementations of SIIB and STOI, the advantage of reducing statistical dependencies between input features is demonstrated. Additionally, the paper presents a new version of SIIB called , which has similar performance to SIIB and HASPI, but takes less time to compute by two orders of magnitude.
Published in IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2018
References in corpus (1)
Cited by in corpus (9)
- A unified convolutional beamformer for simultaneous denoising and dereverberation
- EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy Environments
- Unsupervised Low Latency Speech Enhancement with RT-GCC-NMF
- A Neural-Network Framework for the Design of Individualised Hearing-Loss Compensation
- Speech intelligibility of simulated hearing loss sounds and its prediction using the Gammachirp Envelope Similarity Index (GESI)
- Minimum Processing Near-end Listening Enhancement
- iMetricGAN: Intelligibility Enhancement for Speech-in-Noise using Generative Adversarial Network-based Metric Learning
- Predicting speech intelligibility in older adults for speech enhancement using the Gammachirp Envelope Similarity Index, GESI
- Multi-Metric Optimization using Generative Adversarial Networks for Near-End Speech Intelligibility Enhancement