activity
20212024
most citedMETTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD2024

Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling

Yuepeng Jiang, Tao Li, Fengyu Yang +3

Recent research in zero-shot speech synthesis has made significant progress in speaker similarity. However, current efforts focus on timbre generalization rather than prosody model…

cs.LG2023

Bayesian Domain Invariant Learning via Posterior Generalization of Parameter Distributions

Shiyu Shen, Bin Pan, Tianyang Shi +2

Domain invariant learning aims to learn models that extract invariant features over various training domains, resulting in better generalization to unseen target domains. Recently,…

cs.CR2023

Decision-Dominant Strategic Defense Against Lateral Movement for 5G Zero-Trust Multi-Domain Networks

Tao Li, Yunian Pan, Quanyan Zhu

Multi-domain warfare is a military doctrine that leverages capabilities from different domains, including air, land, sea, space, and cyberspace, to create a highly interconnected b…

eess.AS20231 cited

METTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer

Xinfa Zhu, Yi Lei, Tao Li +4

Previous multilingual text-to-speech (TTS) approaches have considered leveraging monolingual speaker data to enable cross-lingual speech synthesis. However, such data-efficient app…

cs.LG2023

AutoQNN: An End-to-End Framework for Automatically Quantizing Neural Networks

Cheng Gong, Ye Lu, Surong Dai +3

Exploring the expected quantizing scheme with suitable mixed-precision policy is the key point to compress deep neural networks (DNNs) in high efficiency and accuracy. This explora…

eess.AS2021

Multi-speaker Multi-style Text-to-speech Synthesis With Single-speaker Single-style Training Data Scenarios

Qicong Xie, Tao Li, Xinsheng Wang +4

In the existing cross-speaker style transfer task, a source speaker with multi-style recordings is necessary to provide the style for a target speaker. However, it is hard for one…