7 citations · 7 across the 4 of their papers we have counts for
4 papers
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Philip Anastassiou, Jiawei Chen, Jitong Chen +43
We introduce Seed-TTS, a family of large-scale autoregressive text-to-speech (TTS) models capable of generating speech that is virtually indistinguishable from human speech. Seed-T…
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
Yi Yuan, Zhuo Chen, Xubo Liu +6
Contrastive language-audio pretraining~(CLAP) has been developed to align the representations of audio and language, achieving remarkable performance in retrieval and classificatio…
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
Philip Anastassiou, Zhenyu Tang, Kainan Peng +6
We present VoiceShop, a novel speech-to-speech framework that can modify multiple attributes of speech, such as age, gender, accent, and speech style, in a single forward pass whil…
Distinguishing Mechanisms Underlying EMT Tristability
Dongya Jia, Mohit Kumar Jolly, Satyendra C. Tripathi +10
Background: The Epithelial-Mesenchymal Transition (EMT) endows epithelial-looking cells with enhanced migratory ability during embryonic development and tissue repair. EMT can also…