Publications (23)
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
Shu-wen Yang, Ming Tu, Andy T. Liu +5
Speech-to-Speech (S2S) models have shown promising dialogue capabilities, but their ability to handle paralinguistic cues - such as emotion, tone, and speaker attributes - and to r…
Memory Augmented Lookup Dictionary based Language Modeling for Automatic Speech Recognition
Yukun Feng, Ming Tu, Rui Xia +2
Recent studies have shown that using an external Language Model (LM) benefits the end-to-end Automatic Speech Recognition (ASR). However, predicting tokens that appear less frequen…
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
Philip Anastassiou, Zhenyu Tang, Kainan Peng +6
We present VoiceShop, a novel speech-to-speech framework that can modify multiple attributes of speech, such as age, gender, accent, and speech style, in a single forward pass whil…
Investigating the role of L1 in automatic pronunciation evaluation of L2 speech
Ming Tu, Anna Grabek, Julie Liss +1
Automatic pronunciation evaluation plays an important role in pronunciation training and second language education. This field draws heavily on concepts from automatic speech recog…
Language-Universal Phonetic Representation in Multilingual Speech Pretraining for Low-Resource Speech Recognition
Siyuan Feng, Ming Tu, Rui Xia +2
We improve low-resource ASR by integrating the ideas of multilingual training and self-supervised learning. Concretely, we leverage an International Phonetic Alphabet (IPA) multili…
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
Ye Bai, Jingping Chen, Jitong Chen +52
Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific con…
Streaming Voice Conversion Via Intermediate Bottleneck Features And Non-streaming Teacher Guidance
Yuanzhe Chen, Ming Tu, Tang Li +7
Streaming voice conversion (VC) is the task of converting the voice of one person to another in real-time. Previous streaming VC methods use phonetic posteriorgrams (PPGs) extracte…
Multiple instance learning with graph neural networks
Ming Tu, Jing Huang, Xiaodong He +1
Multiple instance learning (MIL) aims to learn the mapping between a bag of instances and the bag-level label. In this paper, we propose a new end-to-end graph neural network (GNN)…
I4U Submission to NIST SRE 2018: Leveraging from a Decade of Shared Experiences
Kong Aik Lee, Ville Hautamaki, Tomi Kinnunen +43
The I4U consortium was established to facilitate a joint entry to NIST speaker recognition evaluations (SRE). The latest edition of such joint submission was in SRE 2018, in which…
Language-universal phonetic encoder for low-resource speech recognition
Siyuan Feng, Ming Tu, Rui Xia +2
Multilingual training is effective in improving low-resource ASR, which may partially be explained by phonetic representation sharing between languages. In end-to-end (E2E) ASR sys…
Multi-hop Reading Comprehension across Multiple Documents by Reasoning over Heterogeneous Graphs
Ming Tu, Guangtao Wang, Jing Huang +3
Multi-hop reading comprehension (RC) across documents poses new challenge over single-document RC because it requires reasoning over multiple documents to reach the final answer. I…
Simulating dysarthric speech for training data augmentation in clinical speech applications
Yishan Jiao, Ming Tu, Visar Berisha +1
Training machine learning algorithms for speech applications requires large, labeled training data sets. This is problematic for clinical applications where obtaining such data is…
Erasing Radio Frequency Fingerprints via Active Adversarial Perturbation
Zhaoyi Lu, Wenchao Xu, Ming Tu +3
Radio Frequency (RF) fingerprinting is to identify a wireless device from its uniqueness of the analog circuitry or hardware imperfections. However, unlike the MAC address which ca…
Efficient Neural Music Generation
Max W. Y. Lam, Qiao Tian, Tang Li +10
Recent progress in music generation has been remarkably advanced by the state-of-the-art MusicLM, which comprises a hierarchy of three LMs, respectively, for semantic, coarse acous…
Graph Sequential Network for Reasoning over Sequences
Ming Tu, Jing Huang, Xiaodong He +1
Recently Graph Neural Network (GNN) has been applied successfully to various NLP tasks that require reasoning, such as multi-hop machine reading comprehension. In this paper, we co…
Cloning one's voice using very limited data in the wild
Dongyang Dai, Yuanzhe Chen, Li Chen +6
With the increasing popularity of speech synthesis products, the industry has put forward more requirements for personalized speech synthesis: (1) How to use low-resource, easily a…
Select, Answer and Explain: Interpretable Multi-hop Reading Comprehension over Multiple Documents
Ming Tu, Kevin Huang, Guangtao Wang +3
Interpretable multi-hop reading comprehension (RC) over multiple documents is a challenging problem because it demands reasoning over multiple information sources and explaining th…
Towards adversarial learning of speaker-invariant representation for speech emotion recognition
Ming Tu, Yun Tang, Jing Huang +2
Speech emotion recognition (SER) has attracted great attention in recent years due to the high demand for emotionally intelligent speech interfaces. Deriving speaker-invariant repr…
Reducing the Model Order of Deep Neural Networks Using Information Theory
Ming Tu, Visar Berisha, Yu Cao +1
Deep neural networks are typically represented by a much larger number of parameters than shallow models, making them prohibitive for small footprint devices. Recent research shows…
FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
Weiting Tan, Andy T. Liu, Ming Tu +3
Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate…
Speaker-invariant Affective Representation Learning via Adversarial Training
Haoqi Li, Ming Tu, Jing Huang +2
Representation learning for speech emotion recognition is challenging due to labeled data sparsity issue and lack of gold standard references. In addition, there is much variabilit…
A Discriminative Acoustic-Prosodic Approach for Measuring Local Entrainment
Megan M. Willi, Stephanie A. Borrie, Tyson S. Barrett +2
Acoustic-prosodic entrainment describes the tendency of humans to align or adapt their speech acoustics to each other in conversation. This alignment of spoken behavior has importa…
Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
Weiting Tan, Xinghua Qu, Ming Tu +4
Effective interactive tool use requires agents to master Tool Integrated Reasoning (TIR): a complex process involving multi-turn planning and long-context dialogue management. To t…