papers

Publications (23)

eess.AS2026

ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction

Shu-wen Yang, Ming Tu, Andy T. Liu +5

Speech-to-Speech (S2S) models have shown promising dialogue capabilities, but their ability to handle paralinguistic cues - such as emotion, tone, and speaker attributes - and to r…

cs.CL2022

Memory Augmented Lookup Dictionary based Language Modeling for Automatic Speech Recognition

Yukun Feng, Ming Tu, Rui Xia +2

Recent studies have shown that using an external Language Model (LM) benefits the end-to-end Automatic Speech Recognition (ASR). However, predicting tokens that appear less frequen…

cs.SD2024

VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing

Philip Anastassiou, Zhenyu Tang, Kainan Peng +6

We present VoiceShop, a novel speech-to-speech framework that can modify multiple attributes of speech, such as age, gender, accent, and speech style, in a single forward pass whil…

eess.AS2018

Investigating the role of L1 in automatic pronunciation evaluation of L2 speech

Ming Tu, Anna Grabek, Julie Liss +1

Automatic pronunciation evaluation plays an important role in pronunciation training and second language education. This field draws heavily on concepts from automatic speech recog…

eess.AS2023

Language-Universal Phonetic Representation in Multilingual Speech Pretraining for Low-Resource Speech Recognition

Siyuan Feng, Ming Tu, Rui Xia +2

We improve low-resource ASR by integrating the ideas of multilingual training and self-supervised learning. Concretely, we leverage an International Phonetic Alphabet (IPA) multili…

eess.AS2024

Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Ye Bai, Jingping Chen, Jitong Chen +52

Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific con…

eess.AS2022

Streaming Voice Conversion Via Intermediate Bottleneck Features And Non-streaming Teacher Guidance

Yuanzhe Chen, Ming Tu, Tang Li +7

Streaming voice conversion (VC) is the task of converting the voice of one person to another in real-time. Previous streaming VC methods use phonetic posteriorgrams (PPGs) extracte…

cs.LG2019

Multiple instance learning with graph neural networks

Ming Tu, Jing Huang, Xiaodong He +1

Multiple instance learning (MIL) aims to learn the mapping between a bag of instances and the bag-level label. In this paper, we propose a new end-to-end graph neural network (GNN)…

eess.AS2019

I4U Submission to NIST SRE 2018: Leveraging from a Decade of Shared Experiences

Kong Aik Lee, Ville Hautamaki, Tomi Kinnunen +43

The I4U consortium was established to facilitate a joint entry to NIST speaker recognition evaluations (SRE). The latest edition of such joint submission was in SRE 2018, in which…

eess.AS2023

Language-universal phonetic encoder for low-resource speech recognition

Siyuan Feng, Ming Tu, Rui Xia +2

Multilingual training is effective in improving low-resource ASR, which may partially be explained by phonetic representation sharing between languages. In end-to-end (E2E) ASR sys…

cs.CL2019

Multi-hop Reading Comprehension across Multiple Documents by Reasoning over Heterogeneous Graphs

Ming Tu, Guangtao Wang, Jing Huang +3

Multi-hop reading comprehension (RC) across documents poses new challenge over single-document RC because it requires reasoning over multiple documents to reach the final answer. I…

eess.AS2018

Simulating dysarthric speech for training data augmentation in clinical speech applications

Yishan Jiao, Ming Tu, Visar Berisha +1

Training machine learning algorithms for speech applications requires large, labeled training data sets. This is problematic for clinical applications where obtaining such data is…

cs.CR2024

Erasing Radio Frequency Fingerprints via Active Adversarial Perturbation

Zhaoyi Lu, Wenchao Xu, Ming Tu +3

Radio Frequency (RF) fingerprinting is to identify a wireless device from its uniqueness of the analog circuitry or hardware imperfections. However, unlike the MAC address which ca…

cs.SD2023

Efficient Neural Music Generation

Max W. Y. Lam, Qiao Tian, Tang Li +10

Recent progress in music generation has been remarkably advanced by the state-of-the-art MusicLM, which comprises a hierarchy of three LMs, respectively, for semantic, coarse acous…

cs.CL2020

Graph Sequential Network for Reasoning over Sequences

Ming Tu, Jing Huang, Xiaodong He +1

Recently Graph Neural Network (GNN) has been applied successfully to various NLP tasks that require reasoning, such as multi-hop machine reading comprehension. In this paper, we co…

eess.AS2021

Cloning one's voice using very limited data in the wild

Dongyang Dai, Yuanzhe Chen, Li Chen +6

With the increasing popularity of speech synthesis products, the industry has put forward more requirements for personalized speech synthesis: (1) How to use low-resource, easily a…

cs.CL2020

Select, Answer and Explain: Interpretable Multi-hop Reading Comprehension over Multiple Documents

Ming Tu, Kevin Huang, Guangtao Wang +3

Interpretable multi-hop reading comprehension (RC) over multiple documents is a challenging problem because it demands reasoning over multiple information sources and explaining th…

eess.AS2019

Towards adversarial learning of speaker-invariant representation for speech emotion recognition

Ming Tu, Yun Tang, Jing Huang +2

Speech emotion recognition (SER) has attracted great attention in recent years due to the high demand for emotionally intelligent speech interfaces. Deriving speaker-invariant repr…

cs.LG2016

Reducing the Model Order of Deep Neural Networks Using Information Theory

Ming Tu, Visar Berisha, Yu Cao +1

Deep neural networks are typically represented by a much larger number of parameters than shallow models, making them prohibitive for small footprint devices. Recent research shows…

cs.CV2026

FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation

Weiting Tan, Andy T. Liu, Ming Tu +3

Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate…

eess.AS2021

Speaker-invariant Affective Representation Learning via Adversarial Training

Haoqi Li, Ming Tu, Jing Huang +2

Representation learning for speech emotion recognition is challenging due to labeled data sparsity issue and lack of gold standard references. In addition, there is much variabilit…

eess.AS2018

A Discriminative Acoustic-Prosodic Approach for Measuring Local Entrainment

Megan M. Willi, Stephanie A. Borrie, Tyson S. Barrett +2

Acoustic-prosodic entrainment describes the tendency of humans to align or adapt their speech acoustics to each other in conversation. This alignment of spoken behavior has importa…

cs.CL2025

Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents

Weiting Tan, Xinghua Qu, Ming Tu +4

Effective interactive tool use requires agents to master Tool Integrated Reasoning (TIR): a complex process involving multi-turn planning and long-context dialogue management. To t…