EffVOC: Low-Delay Efficient Speech Waveform Reconstruction from Spectral Representations Without Phase
arXiv:2609.04226
Abstract
The Griffin-Lim algorithm has been a seminal contribution for phase reconstruction from amplitude spectrograms, however, requiring (infinitely) high algorithmic delay. Its low-delay variant suffers in speech quality. Recent (generative) neural network methods improve on speech quality still at medium to high algorithmic delay, but often they are complex and optimized only for one specific input representation. We build upon an efficient low-delay speech vocoder and propose EffVOC, which supports synthesis of wideband (WB) or fullband (FB) speech from either amplitude spectrum or Mel coefficient inputs. We evaluate both input representations across multiple model sizes in a unified framework and compare against state of the art. Results show that our proposed low-delay (20 ms vs. 32 ms or more) efficient approach marks a new SOTA by achieving top-ranked subjective MOS scores (WB: 4.17/4.15, FB: 4.14/4.11) for amplitude spectrum/Mel representations, very close to ground truth.