2 citations · 2 across the 2 of their papers we have counts for
2 papers
eess.AS2022
Adapter-Based Extension of Multi-Speaker Text-to-Speech Model for New Speakers
Cheng-Ping Hsieh, Subhankar Ghosh, Boris Ginsburg
Fine-tuning is a popular method for adapting text-to-speech (TTS) models to new speakers. However this approach has some challenges. Usually fine-tuning requires several hours of h…
cs.IR2022★ 2 cited
Mr. Right: Multimodal Retrieval on Representation of ImaGe witH Text
Cheng-An Hsieh, Cheng-Ping Hsieh, Pu-Jen Cheng
Multimodal learning is a recent challenge that extends unimodal learning by generalizing its domain to diverse modalities, such as texts, images, or speech. This extension requires…