12 citations · 22 across the 6 of their papers we have counts for
5 papers · 1 filter
S3T: Self-Supervised Pre-training with Swin Transformer for Music Classification
Hang Zhao, Chen Zhang, Belei Zhu +2
In this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive e…
PDAugment: Data Augmentation by Pitch and Duration Adjustments for Automatic Lyrics Transcription
Chen Zhang, Jiaxing Yu, LuChin Chang +4
Automatic lyrics transcription (ALT), which can be regarded as automatic speech recognition (ASR) on singing voice, is an interesting and practical topic in academia and industry.…
DenoiSpeech: Denoising Text to Speech with Frame-Level Noise Modeling
Chen Zhang, Yi Ren, Xu Tan +5
While neural-based text to speech (TTS) models can synthesize natural and intelligible voice, they usually require high-quality speech data, which is costly to collect. In many sce…
FastLR: Non-Autoregressive Lipreading Model with Integrate-and-Fire
Jinglin Liu, Yi Ren, Zhou Zhao +3
Lipreading is an impressive technique and there has been a definite improvement of accuracy in recent years. However, existing methods for lipreading mainly build on autoregressive…
UWSpeech: Speech to Speech Translation for Unwritten Languages
Chen Zhang, Xu Tan, Yi Ren +3
Existing speech to speech translation systems heavily rely on the text of target language: they usually translate source language either to target text and then synthesize target s…