3 papers
cs.SD2024
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
Tan Dat Nguyen, Ji-Hoon Kim, Jeongsoo Choi +4
The goal of this paper is to accelerate codec-based speech synthesis systems with minimum sacrifice to speech quality. We propose an enhanced inference method that allows for flexi…
eess.AS2024
Boosting Unknown-number Speaker Separation with Transformer Decoder-based Attractor
Younglo Lee, Shukjae Choi, Byeong-Yeol Kim +2
We propose a novel speech separation model designed to separate mixtures with an unknown number of speakers. The proposed model stacks 1) a dual-path processing block that can mode…
cs.SD2018
Analysis Acoustic Features for Acoustic Scene Classification and Score fusion of multi-classification systems applied to DCASE 2016 challenge
Sangwook Park, Seongkyu Mun, Younglo Lee +2
This paper describes an acoustic scene classification method which achieved the 4th ranking result in the IEEE AASP challenge of Detection and Classification of Acoustic Scenes and…