activity
20192022
most citedAn Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning

20 citations · 58 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD20228 cited

The Chamber Ensemble Generator: Limitless High-Quality MIR Data via Generative Modeling

Yusong Wu, Josh Gardner, Ethan Manilow +3

Data is the lifeblood of modern machine learning systems, including for those in Music Information Retrieval (MIR). However, MIR has long been mired by small datasets and unreliabl…

eess.AS202120 cited

An Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning

Xinhao Mei, Qiushi Huang, Xubo Liu +10

Automated audio captioning aims to use natural language to describe the content of audio data. This paper presents an audio captioning system with an encoder-decoder architecture,…

eess.AS20204 cited

Peking Opera Synthesis via Duration Informed Attention Network

Yusong Wu, Shengchen Li, Chengzhu Yu +4

Peking Opera has been the most dominant form of Chinese performing art since around 200 years ago. A Peking Opera singer usually exhibits a very strong personal style via introduci…

eess.AS202013 cited

DurIAN-SC: Duration Informed Attention Network based Singing Voice Conversion System

Liqiang Zhang, Chengzhu Yu, Heng Lu +6

Singing voice conversion is converting the timbre in the source singing to the target speaker's voice while keeping singing content the same. However, singing data for target speak…

cs.CL20193 cited

Synthesising Expressiveness in Peking Opera via Duration Informed Attention Network

Yusong Wu, Shengchen Li, Chengzhu Yu +4

This paper presents a method that generates expressive singing voice of Peking opera. The synthesis of expressive opera singing usually requires pitch contours to be extracted as t…

cs.SD201910 cited

Learning Singing From Speech

Liqiang Zhang, Chengzhu Yu, Heng Lu +5

We propose an algorithm that is capable of synthesizing high quality target speaker's singing voice given only their normal speech samples. The proposed algorithm first integrate s…