activity
20172022
most citedDenoiSpeech: Denoising Text to Speech with Frame-Level Noise Modeling

6 citations · 8 across the 3 of their papers we have counts for

collaborators

6 papers

eess.AS20221 cited

S3T: Self-Supervised Pre-training with Swin Transformer for Music Classification

Hang Zhao, Chen Zhang, Belei Zhu +2

In this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive e…

eess.AS20206 cited

DenoiSpeech: Denoising Text to Speech with Frame-Level Noise Modeling

Chen Zhang, Yi Ren, Xu Tan +5

While neural-based text to speech (TTS) models can synthesize natural and intelligible voice, they usually require high-quality speech data, which is costly to collect. In many sce…

eess.AS2020

MusiCoder: A Universal Music-Acoustic Encoder Based on Transformers

Yilun Zhao, Jia Guo

Music annotation has always been one of the critical topics in the field of Music Information Retrieval (MIR). Traditional models use supervised learning for music annotation tasks…

eess.AS2020

UWSpeech: Speech to Speech Translation for Unwritten Languages

Chen Zhang, Xu Tan, Yi Ren +3

Existing speech to speech translation systems heavily rely on the text of target language: they usually translate source language either to target text and then synthesize target s…

cs.CV2019

User independent Emotion Recognition with Residual Signal-Image Network

Guanghao Yin, Shouqian Sun, Hui Zhang +4

User independent emotion recognition with large scale physiological signals is a tough problem. There exist many advanced methods but they are conducted under relatively small data…

cs.CL20171 cited

A Novel Comprehensive Approach for Estimating Concept Semantic Similarity in WordNet

Xiao-gang Zhang, Shou-qian Sun, Ke-jun Zhang

Computation of semantic similarity between concepts is an important foundation for many research works. This paper focuses on IC computing methods and IC measures, which estimate t…