papers

Publications (106)

eess.AS2020

Speaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams

Huirong Huang, Zhiyong Wu, Shiyin Kang +9

Generating 3D speech-driven talking head has received more and more attention in recent years. Recent approaches mainly have following limitations: 1) most speaker-independent meth…

cond-mat.mtrl-sci2022

Multi-state data storage in a two-dimensional stripy antiferromagnet implemented by magnetoelectric effect

Pingfan Gu, Cong Wang, Dan Su +8

A promising approach to the next generation of low-power, functional, and energy-efficient electronics relies on novel materials with coupled magnetic and electric degrees of freed…

cs.CR2020

"Are you home alone?" "Yes" Disclosing Security and Privacy Vulnerabilities in Alexa Skills

Dan Su, Jiqiang Liu, Sencun Zhu +2

The home voice assistants such as Amazon Alexa have become increasingly popular due to many interesting voice-activated services provided through special applications called skills…

cs.CL2019

DurIAN: Duration Informed Attention Network For Multimodal Synthesis

Chengzhu Yu, Heng Lu, Na Hu +9

In this paper, we present a generic and robust multimodal synthesis system that produces highly natural speech and facial expression simultaneously. The key component of this syste…

eess.AS2021

Referee: Towards reference-free cross-speaker style transfer with low-quality data for expressive speech synthesis

Songxiang Liu, Shan Yang, Dan Su +1

Cross-speaker style transfer (CSST) in text-to-speech (TTS) synthesis aims at transferring a speaking style to the synthesised speech in a target speaker's voice. Most previous CSS…

cs.CL2025

Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aaron Blakeman +311

We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 t…