2 citations · 7 across the 10 of their papers we have counts for
14 papers · 1 filter
Improve Bilingual TTS Using Dynamic Language and Phonology Embedding
Fengyu Yang, Jian Luan, Yujun Wang
In most cases, bilingual TTS needs to handle three types of input scripts: first language only, second language only, and second language embedded in the first language. In the lat…
An empirical study of weakly supervised audio tagging embeddings for general audio representations
Heinrich Dinkel, Zhiyong Yan, Yongqing Wang +2
We study the usability of pre-trained weakly supervised audio tagging (AT) models as feature extractors for general audio representations. We mainly analyze the feasibility of tran…
UniKW-AT: Unified Keyword Spotting and Audio Tagging
Heinrich Dinkel, Yongqing Wang, Zhiyong Yan +2
Within the audio research community and the industry, keyword spotting (KWS) and audio tagging (AT) are seen as two distinct tasks and research fields. However, from a technical po…
Pseudo strong labels for large scale weakly supervised audio tagging
Heinrich Dinkel, Zhiyong Yan, Yongqing Wang +2
Large-scale audio tagging datasets inevitably contain imperfect labels, such as clip-wise annotated (temporally weak) tags with no exact on- and offsets, due to a high manual label…
Learning Decoupling Features Through Orthogonality Regularization
Li Wang, Rongzhi Gu, Weiji Zhuang +3
Keyword spotting (KWS) and speaker verification (SV) are two important tasks in speech applications. Research shows that the state-of-art KWS and SV models are trained independentl…
A Separable Temporal Convolution Neural Network with Attention for Small-Footprint Keyword Spotting
Shenghua Hu, Jing Wang, Yujun Wang +2
Keyword spotting (KWS) on mobile devices generally requires a small memory footprint. However, most current models still maintain a large number of parameters in order to ensure go…