2 citations · 3 across the 3 of their papers we have counts for
4 papers · 1 filter
LMCodec: A Low Bitrate Speech Codec With Causal Transformer Models
Teerapat Jenrungrot, Michael Chinen, W. Bastiaan Kleijn +4
We introduce LMCodec, a causal neural speech codec that provides high quality audio at very low bitrates. The backbone of the system is a causal convolutional codec that encodes au…
Using Rater and System Metadata to Explain Variance in the VoiceMOS Challenge 2022 Dataset
Michael Chinen, Jan Skoglund, Chandan K A Reddy +2
Non-reference speech quality models are important for a growing number of applications. The VoiceMOS 2022 challenge provided a dataset of synthetic voice conversion and text-to-spe…
Differentiable Consistency Constraints for Improved Deep Speech Enhancement
Scott Wisdom, John R. Hershey, Kevin Wilson +4
In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement system…
Exploring Tradeoffs in Models for Low-latency Speech Enhancement
Kevin Wilson, Michael Chinen, Jeremy Thorpe +5
We explore a variety of neural networks configurations for one- and two-channel spectrogram-mask-based speech enhancement. Our best model improves on previous state-of-the-art perf…