papers

Publications (220)

cs.CL2025

Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions

Lingwei Meng, Shujie Hu, Jiawen Kang +6

Recent advancements in large language models (LLMs) have revolutionized various domains, bringing significant progress and new opportunities. Despite progress in speech-related tas…

cs.CL2025

Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching

Xiaoying Zhang, Baolin Peng, Ye Tian +4

Large language models (LLMs) often struggle to provide up-to-date information due to their one-time training and the constantly evolving nature of the world. To keep LLMs current,…

cs.SD2025

AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction

Wenxuan Wu, Xueyuan Chen, Shuai Wang +5

Audio-Visual Target Speaker Extraction (AV-TSE) aims to mimic the human ability to enhance auditory perception using visual cues. Although numerous models have been proposed recent…

cs.SD2023

InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Dongchao Yang, Songxiang Liu, Rongjie Huang +2

Expressive text-to-speech (TTS) aims to synthesize different speaking style speech according to human's demands. Nowadays, there are two common ways to control speaking styles: (1)…

cs.SD2024

Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder

Haohan Guo, Fenglong Xie, Dongchao Yang +3

VQ-VAE, as a mainstream approach of speech tokenizer, has been troubled by ``index collapse'', where only a small number of codewords are activated in large codebooks. This work pr…

eess.AS2020

Transferring Source Style in Non-Parallel Voice Conversion

Songxiang Liu, Yuewen Cao, Shiyin Kang +5

Voice conversion (VC) techniques aim to modify speaker identity of an utterance while preserving the underlying linguistic information. Most VC approaches ignore modeling of the sp…

cs.SD2024

Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy

Wenxuan Wu, Xueyuan Chen, Xixin Wu +2

Audio-visual target speech extraction (AV-TSE) is one of the enabling technologies in robotics and many audio-visual applications. One of the challenges of AV-TSE is how to effecti…

cs.SD2022

Spoofing-Aware Speaker Verification by Multi-Level Fusion

Haibin Wu, Lingwei Meng, Jiawen Kang +5

Recently, many novel techniques have been introduced to deal with spoofing attacks, and achieve promising countermeasure (CM) performances. However, these works only take the stand…

cs.CL2021

Open Intent Discovery through Unsupervised Semantic Clustering and Dependency Parsing

Pengfei Liu, Youzhang Ning, King Keung Wu +2

Intent understanding plays an important role in dialog systems, and is typically formulated as a supervised learning problem. However, it is challenging and time-consuming to desig…

cs.SD2024

Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation

Haohan Guo, Fenglong Xie, Dongchao Yang +2

The neural codec language model (CLM) has demonstrated remarkable performance in text-to-speech (TTS) synthesis. However, troubled by ``recency bias", CLM lacks sufficient attentio…

cs.SD2024

Improving the Adversarial Robustness for Speaker Verification by Self-Supervised Learning

Haibin Wu, Xu Li, Andy T. Liu +3

Previous works have shown that automatic speaker verification (ASV) is seriously vulnerable to malicious spoofing attacks, such as replay, synthetic speech, and recently emerged ad…

eess.AS2020

Audio-visual Multi-channel Recognition of Overlapped Speech

Jianwei Yu, Bo Wu, Rongzhi Gu +7

Automatic speech recognition (ASR) of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data are widely used in state-of-the-…

eess.AS2023

Consistent and Relevant: Rethink the Query Embedding in General Sound Separation

Yuanyuan Wang, Hangting Chen, Dongchao Yang +4

The query-based audio separation usually employs specific queries to extract target sources from a mixture of audio signals. Currently, most query-based separation models need addi…

cs.SD2022

Neural Architecture Search for Speech Emotion Recognition

Xixin Wu, Shoukang Hu, Zhiyong Wu +2

Deep neural networks have brought significant advancements to speech emotion recognition (SER). However, the architecture design in SER is mainly based on expert knowledge and empi…

cs.CL2022

User Satisfaction Estimation with Sequential Dialogue Act Modeling in Goal-oriented Conversational Systems

Yang Deng, Wenxuan Zhang, Wai Lam +2

User Satisfaction Estimation (USE) is an important yet challenging task in goal-oriented conversational systems. Whether the user is satisfied with the system largely depends on th…

cs.LG2023

Decision Support System for Chronic Diseases Based on Drug-Drug Interactions

Tian Bian, Yuli Jiang, Jia Li +6

Many patients with chronic diseases resort to multiple medications to relieve various symptoms, which raises concerns about the safety of multiple medication use, as severe drug-dr…

cs.SD2022

FullSubNet+: Channel Attention FullSubNet with Complex Spectrograms for Speech Enhancement

Jun Chen, Zilin Wang, Deyi Tuo +3

Previously proposed FullSubNet has achieved outstanding performance in Deep Noise Suppression (DNS) Challenge and attracted much attention. However, it still encounters issues such…

eess.AS2020

Improved End-to-End Dysarthric Speech Recognition via Meta-learning Based Model Re-initialization

Disong Wang, Jianwei Yu, Xixin Wu +3

Dysarthric speech recognition is a challenging task as dysarthric data is limited and its acoustics deviate significantly from normal speech. Model-based speaker adaptation is a pr…

eess.AS2022

Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using -VAE

Hui Lu, Disong Wang, Xixin Wu +3

We propose an unsupervised learning method to disentangle speech into content representation and speaker identity representation. We apply this method to the challenging one-shot c…

eess.AS2022

Speaker Adaptation Using Spectro-Temporal Deep Features for Dysarthric and Elderly Speech Recognition

Mengzhe Geng, Xurong Xie, Zi Ye +5

Despite the rapid progress of automatic speech recognition (ASR) technologies targeting normal speech in recent decades, accurate recognition of dysarthric and elderly speech remai…

cs.SD2024

Towards Effective and Efficient Non-autoregressive Decoding Using Block-based Attention Mask

Tianzi Wang, Xurong Xie, Zhaoqing Li +9

This paper proposes a novel non-autoregressive (NAR) block-based Attention Mask Decoder (AMD) that flexibly balances performance-efficiency trade-offs for Conformer ASR systems. AM…

cs.CL2023

Interpretable Unified Language Checking

Tianhua Zhang, Hongyin Luo, Yung-Sung Chuang +7

Despite recent concerns about undesirable behaviors generated by large language models (LLMs), including non-factual, biased, and hateful language, we find LLMs are inherent multi-…

cs.SD2023

StyleSpeech: Self-supervised Style Enhancing with VQ-VAE-based Pre-training for Expressive Audiobook Speech Synthesis

Xueyuan Chen, Xi Wang, Shaofei Zhang +4

The expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data. To address these is…

cs.SD2023

GTN-Bailando: Genre Consistent Long-Term 3D Dance Generation based on Pre-trained Genre Token Network

Haolin Zhuang, Shun Lei, Long Xiao +6

Music-driven 3D dance generation has become an intensive research topic in recent years with great potential for real-world applications. Most existing methods lack the considerati…

cs.CL2022

Toward Self-learning End-to-End Task-Oriented Dialog Systems

Xiaoying Zhang, Baolin Peng, Jianfeng Gao +1

End-to-end task bots are typically learned over a static and usually limited-size corpus. However, when deployed in dynamic, changing, and open environments to interact with users,…

cs.SD2023

Towards Improving the Expressiveness of Singing Voice Synthesis with BERT Derived Semantic Information

Shaohuan Zhou, Shun Lei, Weiya You +5

This paper presents an end-to-end high-quality singing voice synthesis (SVS) system that uses bidirectional encoder representation from Transformers (BERT) derived semantic embeddi…

cs.AI2024

Ontology-grounded Automatic Knowledge Graph Construction by LLM under Wikidata schema

Xiaohan Feng, Xixin Wu, Helen Meng

We propose an ontology-grounded approach to Knowledge Graph (KG) construction using Large Language Models (LLMs) on a knowledge base. An ontology is authored by generating Competen…

cs.SD2024

UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner

Dongchao Yang, Haohan Guo, Yuanyuan Wang +5

The Large Language models (LLMs) have demonstrated supreme capabilities in text understanding and generation, but cannot be directly applied to cross-modal tasks without fine-tunin…

eess.AS2021

Adversarial Data Augmentation for Disordered Speech Recognition

Zengrui Jin, Mengzhe Geng, Xurong Xie +4

Automatic recognition of disordered speech remains a highly challenging task to date. The underlying neuro-motor conditions, often compounded with co-occurring physical disabilitie…

cs.CL2026

TriggerBench: Investigating Prospective Memory for Large Language Models

Tianhua Zhang, Xinjiang Wang, Qianxi Zhang +6

While Large Language Models (LLMs) are increasingly deployed in long interactions, existing evaluations focus predominantly on retrospective memory (RM) via explicit queries. Prosp…

cs.CL2026

TreePS-RAG: Tree-based Process Supervision for Reinforcement Learning in Agentic RAG

Tianhua Zhang, Kun Li, Junan Li +5

Agentic retrieval-augmented generation (RAG) formulates question answering as a multi-step interaction between reasoning and information retrieval, and has recently been advanced b…

cs.SD2023

A Sidecar Separator Can Convert a Single-Talker Speech Recognition System to a Multi-Talker One

Lingwei Meng, Jiawen Kang, Mingyu Cui +3

Although automatic speech recognition (ASR) can perform well in common non-overlapping environments, sustaining performance in multi-talker overlapping speech recognition remains c…

cs.SD2022

Adversarial Sample Detection for Speaker Verification by Neural Vocoders

Haibin Wu, Po-chun Hsu, Ji Gao +6

Automatic speaker verification (ASV), one of the most important technology for biometric identification, has been widely adopted in security-critical applications. However, ASV is…

eess.AS2020

Unsupervised Cross-Lingual Speech Emotion Recognition Using DomainAdversarial Neural Network

Xiong Cai, Zhiyong Wu, Kuo Zhong +3

By using deep learning approaches, Speech Emotion Recog-nition (SER) on a single domain has achieved many excellentresults. However, cross-domain SER is still a challenging taskdue…

cs.SD2024

SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models

Dongchao Yang, Rongjie Huang, Yuanyuan Wang +5

Scaling Text-to-speech (TTS) to large-scale datasets has been demonstrated as an effective method for improving the diversity and naturalness of synthesized speech. At the high lev…

cs.CL2022

Robust Unsupervised Cross-Lingual Word Embedding using Domain Flow Interpolation

Liping Tang, Zhen Li, Zhiquan Luo +1

This paper investigates an unsupervised approach towards deriving a universal, cross-lingual word embedding space, where words with similar semantics from different languages are c…

cs.SD2021

Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input

Xingchen Song, Zhiyong Wu, Yiheng Huang +3

Non-autoregressive (NAR) transformer models have achieved significantly inference speedup but at the cost of inferior accuracy compared to autoregressive (AR) models in automatic s…

cs.SD2023

QS-TTS: Towards Semi-Supervised Text-to-Speech Synthesis via Vector-Quantized Self-Supervised Speech Representation Learning

Haohan Guo, Fenglong Xie, Jiawen Kang +3

This paper proposes a novel semi-supervised TTS framework, QS-TTS, to improve TTS quality with lower supervised data requirements via Vector-Quantized Self-Supervised Speech Repres…

eess.AS2025

learning discriminative features from spectrograms using center loss for speech emotion recognition

Dongyang Dai, Zhiyong Wu, Runnan Li +3

Identifying the emotional state from speech is essential for the natural interaction of the machine with the speaker. However, extracting effective features for emotion recognition…

cs.SD2023

Enhancing the vocal range of single-speaker singing voice synthesis with melody-unsupervised pre-training

Shaohuan Zhou, Xu Li, Zhiyong Wu +2

The single-speaker singing voice synthesis (SVS) usually underperforms at pitch values that are out of the singer's vocal range or associated with limited training samples. Based o…

eess.AS2022

Confidence Score Based Conformer Speaker Adaptation for Speech Recognition

Jiajun Deng, Xurong Xie, Tianzi Wang +7

A key challenge for automatic speech recognition (ASR) systems is to model the speaker level variability. In this paper, compact speaker dependent learning hidden unit contribution…

eess.AS2020

Bayesian x-vector: Bayesian Neural Network based x-vector System for Speaker Verification

Xu Li, Jinghua Zhong, Jianwei Yu +4

Speaker verification systems usually suffer from the mismatch problem between training and evaluation data, such as speaker population mismatch, the channel and environment variati…

cs.SD2026

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions

Junjie Li, Wenxuan Wu, Shuai Wang +4

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate a target speaker's voice from multi-speaker environments by leveraging visual cues as guidance. However, the perform…

cs.AI2024

Improving Grapheme-to-Phoneme Conversion through In-Context Knowledge Retrieval with Large Language Models

Dongrui Han, Mingyu Cui, Jiawen Kang +3

Grapheme-to-phoneme (G2P) conversion is a crucial step in Text-to-Speech (TTS) systems, responsible for mapping grapheme to corresponding phonetic representations. However, it face…

cs.SD2024

SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis

Haohan Guo, Fenglong Xie, Kun Xie +4

The long speech sequence has been troubling language models (LM) based TTS approaches in terms of modeling complexity and efficiency. This work proposes SoCodec, a semantic-ordered…

cs.SD2022

Content-Dependent Fine-Grained Speaker Embedding for Zero-Shot Speaker Adaptation in Text-to-Speech Synthesis

Yixuan Zhou, Changhe Song, Xiang Li +5

Zero-shot speaker adaptation aims to clone an unseen speaker's voice without any adaptation time and parameters. Previous researches usually use a speaker encoder to extract a glob…

cs.SD2024

Multi-view MidiVAE: Fusing Track- and Bar-view Representations for Long Multi-track Symbolic Music Generation

Zhiwei Lin, Jun Chen, Boshi Tang +7

Variational Autoencoders (VAEs) constitute a crucial component of neural symbolic music generation, among which some works have yielded outstanding results and attracted considerab…

cs.CL2021

Unstructured Knowledge Access in Task-oriented Dialog Modeling using Language Inference, Knowledge Retrieval and Knowledge-Integrative Response Generation

Mudit Chaudhary, Borislav Dzodzo, Sida Huang +10

Dialog systems enriched with external knowledge can handle user queries that are outside the scope of the supporting databases/APIs. In this paper, we follow the baseline provided…

cs.SD2022

Towards Expressive Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis

Shun Lei, Yixuan Zhou, Liyang Chen +3

Previous works on expressive speech synthesis mainly focus on current sentence. The context in adjacent sentences is neglected, resulting in inflexible speaking style for the same…

cs.CL2026

Lamer-SSL: Layer-aware Mixture of LoRA Experts for Continual Multilingual Expansion of Self-supervised Models without Forgetting

Jing Xu, Minglin Wu, Xueyuan Chen +2

Despite their impressive performance, self-supervised speech models often struggle to generalize to new languages and tend to forget previously acquired knowledge during continual…

cs.SD2026

UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization

Dongchao Yang, Yuanyuan Wang, Dading Chong +3

We study two foundational problems in audio language models: (1) how to design an audio tokenizer that can serve as an intermediate representation for both understanding and genera…

cs.SD2022

A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS

Haohan Guo, Fenglong Xie, Frank K. Soong +2

We propose a Multi-Stage, Multi-Codebook (MSMC) approach to high-performance neural TTS synthesis. A vector-quantized, variational autoencoder (VQ-VAE) based feature analyzer is us…

cs.SD2021

Mixed Precision DNN Qunatization for Overlapped Speech Separation and Recognition

Junhao Xu, Jianwei Yu, Xunying Liu +1

Recognition of overlapped speech has been a highly challenging task to date. State-of-the-art multi-channel speech separation system are becoming increasingly complex and expensive…

cs.SD2022

Towards Cross-speaker Reading Style Transfer on Audiobook Dataset

Xiang Li, Changhe Song, Xianhao Wei +3

Cross-speaker style transfer aims to extract the speech style of the given reference speech, which can be reproduced in the timbre of arbitrary target speakers. Existing methods on…

eess.AS2021

FastSVC: Fast Cross-Domain Singing Voice Conversion with Feature-wise Linear Modulation

Songxiang Liu, Yuewen Cao, Na Hu +2

This paper presents FastSVC, a light-weight cross-domain singing voice conversion (SVC) system, which can achieve high conversion performance, with inference speed 4x faster than r…

cs.CL2024

Not All Errors Are Equal: Investigation of Speech Recognition Errors in Alzheimer's Disease Detection

Jiawen Kang, Junan Li, Jinchao Li +2

Automatic Speech Recognition (ASR) plays an important role in speech-based automatic detection of Alzheimer's disease (AD). However, recognition errors could propagate downstream,…

cs.SD2023

The defender's perspective on automatic speaker verification: An overview

Haibin Wu, Jiawen Kang, Lingwei Meng +2

Automatic speaker verification (ASV) plays a critical role in security-sensitive environments. Regrettably, the reliability of ASV has been undermined by the emergence of spoofing…

cs.CL2022

Enhancing Word-Level Semantic Representation via Dependency Structure for Expressive Text-to-Speech Synthesis

Yixuan Zhou, Changhe Song, Jingbei Li +4

Exploiting rich linguistic information in raw text is crucial for expressive text-to-speech (TTS). As large scale pre-trained text representation develops, bidirectional encoder re…

cs.SD2026

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

Dingdong Wang, Shujie Liu, Tianhua Zhang +3

Emotional information in speech plays a unique role in multimodal perception. However, current Speech Large Language Models (SpeechLLMs), similar to conventional speech emotion rec…

cs.SD2025

ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction

Wenxuan Wu, Shuai Wang, Xixin Wu +2

Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic…

cs.CV2025

Will Pre-Training Ever End? A First Step Toward Next-Generation Foundation MLLMs via Self-Improving Systematic Cognition

Xiaoying Zhang, Da Peng, Yipeng Zhang +7

Recent progress in (multimodal) large language models ((M)LLMs) has shifted focus from pre-training to inference-time computation and post-training optimization, largely due to con…

cs.SD2024

Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System

Lingwei Meng, Jiawen Kang, Yuejiao Wang +4

Multi-talker speech recognition and target-talker speech recognition, both involve transcription in multi-talker contexts, remain significant challenges. However, existing methods…

cs.SD2024

Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition

Guinan Li, Jiajun Deng, Youjun Chen +8

This paper proposes joint speaker feature learning methods for zero-shot adaptation of audio-visual multichannel speech separation and recognition systems. xVector and ECAPA-TDNN s…

cs.CL2021

Mixed Precision of Quantization of Transformer Language Models for Speech Recognition

Junhao Xu, Shoukang Hu, Jianwei Yu +2

State-of-the-art neural language models represented by Transformers are becoming increasingly complex and expensive for practical applications. Low-bit deep neural network quantiza…

cs.CL2020

Syntactic representation learning for neural network based TTS with syntactic parse tree traversal

Changhe Song, Jingbei Li, Yixuan Zhou +2

Syntactic structure of a sentence text is correlated with the prosodic structure of the speech that is crucial for improving the prosody and naturalness of a text-to-speech (TTS) s…

cs.CL2025

Naturalistic Language-related Movie-Watching fMRI Task for Detecting Neurocognitive Decline and Disorder

Yuejiao Wang, Xianmin Gong, Xixin Wu +4

Early detection is crucial for timely intervention aimed at preventing and slowing the progression of neurocognitive disorder (NCD), a common and significant health problem among t…

cs.SD2022

Towards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis

Shun Lei, Yixuan Zhou, Liyang Chen +4

Previous works on expressive speech synthesis focus on modelling the mono-scale style embedding from the current sentence or context, but the multi-scale nature of speaking style i…

cs.CL2022

An End-to-end Chinese Text Normalization Model based on Rule-guided Flat-Lattice Transformer

Wenlin Dai, Changhe Song, Xiang Li +4

Text normalization, defined as a procedure transforming non standard words to spoken-form words, is crucial to the intelligibility of synthesized speech in text-to-speech system. R…

cs.MM2024

Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information

Qiaochu Huang, Xu He, Boshi Tang +6

Dance generation, as a branch of human motion generation, has attracted increasing attention. Recently, a few works attempt to enhance dance expressiveness, which includes genre ma…

cs.SD2024

Homogeneous Speaker Features for On-the-Fly Dysarthric and Elderly Speaker Adaptation

Mengzhe Geng, Xurong Xie, Jiajun Deng +7

The application of data-intensive automatic speech recognition (ASR) technologies to dysarthric and elderly adult speech is confronted by their mismatch against healthy and nonaged…

cs.SD2022

Characterizing the adversarial vulnerability of speech self-supervised learning

Haibin Wu, Bo Zheng, Xu Li +3

A leaderboard named Speech processing Universal PERformance Benchmark (SUPERB), which aims at benchmarking the performance of a shared self-supervised learning (SSL) speech model a…

cs.SD2024

SongCreator: Lyrics-based Universal Song Generation

Shun Lei, Yixuan Zhou, Boshi Tang +7

Music is an integral part of human culture, embodying human intelligence and creativity, of which songs compose an essential part. While various aspects of song generation have bee…

cs.SD2022

Tackling Spoofing-Aware Speaker Verification with Multi-Model Fusion

Haibin Wu, Jiawen Kang, Lingwei Meng +5

Recent years have witnessed the extraordinary development of automatic speaker verification (ASV). However, previous works show that state-of-the-art ASV models are seriously vulne…

cs.MA2026

Agentic Cognitive Profiling: Realigning Automated Alzheimer's Disease Detection with Clinical Construct Validity

Jiawen Kang, Kun Li, Dongrui Han +5

Automated Alzheimer's Disease (AD) screening has predominantly followed the inductive paradigm of pattern recognition, which directly maps the input signal to the outcome label. Th…

eess.AS2024

Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition

Shujie Hu, Xurong Xie, Mengzhe Geng +8

Self-supervised learning (SSL) based speech foundation models have been applied to a wide range of ASR tasks. However, their application to dysarthric and elderly speech via data-i…

cs.CL2025

RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning

Kun Li, Yunxiang Li, Tianhua Zhang +4

Robust evaluation is critical for deploying trustworthy retrieval-augmented generation (RAG) systems. However, current LLM-based evaluation frameworks predominantly rely on directl…

cs.SD2023

Improving Mandarin Prosodic Structure Prediction with Multi-level Contextual Information

Jie Chen, Changhe Song, Deyi Tuo +4

For text-to-speech (TTS) synthesis, prosodic structure prediction (PSP) plays an important role in producing natural and intelligible speech. Although inter-utterance linguistic in…

eess.AS2025

Spectral-Aware Low-Rank Adaptation for Speaker Verification

Zhe Li, Man-wai Mak, Mert Pilanci +2

Previous research has shown that the principal singular vectors of a pre-trained model's weight matrices capture critical knowledge. In contrast, those associated with small singul…

cs.CL2022

Cross-lingual Word Embeddings in Hyperbolic Space

Chandni Saxena, Mudit Chaudhary, Helen Meng

Cross-lingual word embeddings can be applied to several natural language processing applications across multiple languages. Unlike prior works that use word embeddings based on the…

cs.CL2020

Deep segmental phonetic posterior-grams based discovery of non-categories in L2 English speech

Xu Li, Xixin Wu, Xunying Liu +1

Second language (L2) speech is often labeled with the native, phone categories. However, in many cases, it is difficult to decide on a categorical phone that an L2 segment belongs…

cs.CL2024

A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models

Dingdong Wang, Mingyu Cui, Dongchao Yang +2

With the rise of Speech Large Language Models (Speech LLMs), there has been growing interest in discrete speech tokens for their ability to integrate with text-based tokens seamles…

cs.CL2024

Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational Answers

Tianhua Zhang, Kun Li, Hongyin Luo +3

Query rewriting is a crucial technique for passage retrieval in open-domain conversational question answering (CQA). It decontexualizes conversational queries into self-contained q…

cs.SD2023

Towards Spontaneous Style Modeling with Semi-supervised Pre-training for Conversational Text-to-Speech Synthesis

Weiqin Li, Shun Lei, Qiaochu Huang +4

The spontaneous behavior that often occurs in conversations makes speech more human-like compared to reading-style. However, synthesizing spontaneous-style speech is challenging du…

cs.SD2022

Spectro-Temporal Deep Features for Disordered Speech Assessment and Recognition

Mengzhe Geng, Shansong Liu, Jianwei Yu +6

Automatic recognition of disordered speech remains a highly challenging task to date. Sources of variability commonly found in normal speech including accent, age or gender, when f…

cs.CY2023

Learning Analytics from Spoken Discussion Dialogs in Flipped Classroom

Hang Su, Borislav Dzodzo, Changlun Li +5

The flipped classroom is a new pedagogical strategy that has been gaining increasing importance recently. Spoken discussion dialog commonly occurs in flipped classroom, which embed…

eess.AS2020

Speaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams

Huirong Huang, Zhiyong Wu, Shiyin Kang +9

Generating 3D speech-driven talking head has received more and more attention in recent years. Recent approaches mainly have following limitations: 1) most speaker-independent meth…

cs.SD2023

MSStyleTTS: Multi-Scale Style Modeling with Hierarchical Context Information for Expressive Speech Synthesis

Shun Lei, Yixuan Zhou, Liyang Chen +4

Expressive speech synthesis is crucial for many human-computer interaction scenarios, such as audiobooks, podcasts, and voice assistants. Previous works focus on predicting the sty…

eess.AS2022

Two-pass Decoding and Cross-adaptation Based System Combination of End-to-end Conformer and Hybrid TDNN ASR Systems

Mingyu Cui, Jiajun Deng, Shoukang Hu +7

Fundamental modelling differences between hybrid and end-to-end (E2E) automatic speech recognition (ASR) systems create large diversity and complementarity among them. This paper i…

cs.CL2025

Autoregressive Speech Synthesis without Vector Quantization

Lingwei Meng, Long Zhou, Shujie Liu +9

We present MELLE, a novel continuous-valued token based language modeling approach for text-to-speech synthesis (TTS). MELLE autoregressively generates continuous mel-spectrogram f…

eess.AS2025

Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC

Jiawen Kang, Lingwei Meng, Mingyu Cui +4

Multi-talker speech recognition (MTASR) faces unique challenges in disentangling and transcribing overlapping speech. To address these challenges, this paper investigates the role…

eess.AS2022

VCVTS: Multi-speaker Video-to-Speech synthesis via cross-modal knowledge transfer from voice conversion

Disong Wang, Shan Yang, Dan Su +3

Though significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speec…

cs.SD2024

UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization

Yuejiao Wang, Xixin Wu, Disong Wang +2

Dysarthric speech reconstruction (DSR) systems aim to automatically convert dysarthric speech into normal-sounding speech. The technology eases communication with speakers affected…

eess.AS2020

Defense against adversarial attacks on spoofing countermeasures of ASV

Haibin Wu, Songxiang Liu, Helen Meng +1

Various forefront countermeasure methods for automatic speaker verification (ASV) with considerable performance in anti-spoofing are proposed in the ASVspoof 2019 challenge. Howeve…

cs.SD2022

MFA-Conformer: Multi-scale Feature Aggregation Conformer for Automatic Speaker Verification

Yang Zhang, Zhiqiang Lv, Haibin Wu +5

In this paper, we present Multi-scale Feature Aggregation Conformer (MFA-Conformer), an easy-to-implement, simple but effective backbone for automatic speaker verification based on…

cs.SD2023

SememeASR: Boosting Performance of End-to-End Speech Recognition against Domain and Long-Tailed Data Shift with Sememe Semantic Knowledge

Jiaxu Zhu, Changhe Song, Zhiyong Wu +1

Recently, excellent progress has been made in speech recognition. However, pure data-driven approaches have struggled to solve the problem in domain-mismatch and long-tailed data.…

cs.SD2025

ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling

Dongchao Yang, Songxiang Liu, Haohan Guo +9

Recent advancements in audio language models have underscored the pivotal role of audio tokenization, which converts audio signals into discrete tokens, thereby facilitating the ap…

eess.AS2024

SCNet: Sparse Compression Network for Music Source Separation

Weinan Tong, Jiaxu Zhu, Jun Chen +5

Deep learning-based methods have made significant achievements in music source separation. However, obtaining good results while maintaining a low model complexity remains challeng…

eess.AS2023

Hyper-parameter Adaptation of Conformer ASR Systems for Elderly and Dysarthric Speech Recognition

Tianzi Wang, Shoukang Hu, Jiajun Deng +5

Automatic recognition of disordered and elderly speech remains highly challenging tasks to date due to data scarcity. Parameter fine-tuning is often used to exploit the large quant…

eess.AS2021

Unsupervised Domain Adaptation for Dysarthric Speech Detection via Domain Adversarial Training and Mutual Information Minimization

Disong Wang, Liqun Deng, Yu Ting Yeung +3

Dysarthric speech detection (DSD) systems aim to detect characteristics of the neuromotor disorder from speech. Such systems are particularly susceptible to domain mismatch where t…

eess.AS2020

Multi-Target Emotional Voice Conversion With Neural Vocoders

Songxiang Liu, Yuewen Cao, Helen Meng

Emotional voice conversion (EVC) is one way to generate expressive synthetic speech. Previous approaches mainly focused on modeling one-to-one mapping, i.e., conversion from one em…