activity
20202022
most citedSongDriver: Real-time Music Accompaniment Generation without Logical Latency nor Exposure Bias

12 citations · 22 across the 6 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS20221 cited

S3T: Self-Supervised Pre-training with Swin Transformer for Music Classification

Hang Zhao, Chen Zhang, Belei Zhu +2

In this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive e…

eess.AS20213 cited

PDAugment: Data Augmentation by Pitch and Duration Adjustments for Automatic Lyrics Transcription

Chen Zhang, Jiaxing Yu, LuChin Chang +4

Automatic lyrics transcription (ALT), which can be regarded as automatic speech recognition (ASR) on singing voice, is an interesting and practical topic in academia and industry.…

eess.AS20206 cited

DenoiSpeech: Denoising Text to Speech with Frame-Level Noise Modeling

Chen Zhang, Yi Ren, Xu Tan +5

While neural-based text to speech (TTS) models can synthesize natural and intelligible voice, they usually require high-quality speech data, which is costly to collect. In many sce…

eess.AS2020

FastLR: Non-Autoregressive Lipreading Model with Integrate-and-Fire

Jinglin Liu, Yi Ren, Zhou Zhao +3

Lipreading is an impressive technique and there has been a definite improvement of accuracy in recent years. However, existing methods for lipreading mainly build on autoregressive…

eess.AS2020

UWSpeech: Speech to Speech Translation for Unwritten Languages

Chen Zhang, Xu Tan, Yi Ren +3

Existing speech to speech translation systems heavily rely on the text of target language: they usually translate source language either to target text and then synthesize target s…