Publications (106)
Speaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams
Huirong Huang, Zhiyong Wu, Shiyin Kang +9
Generating 3D speech-driven talking head has received more and more attention in recent years. Recent approaches mainly have following limitations: 1) most speaker-independent meth…
Multi-state data storage in a two-dimensional stripy antiferromagnet implemented by magnetoelectric effect
Pingfan Gu, Cong Wang, Dan Su +8
A promising approach to the next generation of low-power, functional, and energy-efficient electronics relies on novel materials with coupled magnetic and electric degrees of freed…
"Are you home alone?" "Yes" Disclosing Security and Privacy Vulnerabilities in Alexa Skills
Dan Su, Jiqiang Liu, Sencun Zhu +2
The home voice assistants such as Amazon Alexa have become increasingly popular due to many interesting voice-activated services provided through special applications called skills…
DurIAN: Duration Informed Attention Network For Multimodal Synthesis
Chengzhu Yu, Heng Lu, Na Hu +9
In this paper, we present a generic and robust multimodal synthesis system that produces highly natural speech and facial expression simultaneously. The key component of this syste…
Referee: Towards reference-free cross-speaker style transfer with low-quality data for expressive speech synthesis
Songxiang Liu, Shan Yang, Dan Su +1
Cross-speaker style transfer (CSST) in text-to-speech (TTS) synthesis aims at transferring a speaking style to the synthesised speech in a target speaker's voice. Most previous CSS…
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA, :, Aaron Blakeman +311
We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 t…