activity
20172023
most citedPsychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio Coding

30 citations · 42 across the 19 of their papers we have counts for

collaborators
Showing cs.SDShow all

8 papers · 1 filter

cs.SD202030 cited

Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio Coding

Kai Zhen, Mi Suk Lee, Jongmo Sung +2

Conventional audio coding technologies commonly leverage human perception of sound, or psychoacoustics, to reduce the bitrate while preserving the perceptual quality of the decoded…

cs.SD20203 cited

Deep Autotuner: a Pitch Correcting Network for Singing Performances

Sanna Wager, George Tzanetakis, Cheng-i Wang +1

We introduce a data-driven approach to automatic pitch correction of solo singing performances. The proposed approach predicts note-wise pitch shifts from the relationship between…

cs.SD2019

A Dual-Staged Context Aggregation Method Towards Efficient End-To-End Speech Enhancement

Kai Zhen, Mi Suk Lee, Minje Kim

In speech enhancement, an end-to-end deep neural network converts a noisy speech signal to a clean speech directly in time domain without time-frequency transformation or mask esti…

cs.SD20192 cited

Deep Autotuner: A Data-Driven Approach to Natural-Sounding Pitch Correction for Singing Voice in Karaoke Performances

Sanna Wager, George Tzanetakis, Cheng-i Wang +3

We describe a machine-learning approach to pitch correcting a solo singing performance in a karaoke setting, where the solo voice and accompaniment are on separate tracks. The prop…

cs.SD2018

A Data-Driven Approach to Smooth Pitch Correction for Singing Voice in Pop Music

Sanna Wager, Lijiang Guo, Aswin Sivaraman +1

In this paper, we present a machine-learning approach to pitch correction for voice in a karaoke setting, where the vocals and accompaniment are on separate tracks and time-aligned…

cs.SD20183 cited

On Psychoacoustically Weighted Cost Functions Towards Resource-Efficient Deep Neural Networks for Speech Denoising

Kai Zhen, Aswin Sivaraman, Jongmo Sung +1

We present a psychoacoustically enhanced cost function to balance network complexity and perceptual performance of deep neural networks for speech denoising. While training the net…