papers

Publications (85)

astro-ph.IM2018

Harmonic Summing Improves Pulsar Detection Sensitivity: A Probability Analysis

Meng Yu

Practical application of the harmonic summing technique in the power-spectrum analysis for searching pulsars has exhibited the technique's effectiveness. In this paper, theoretical…

cs.CV2026

Elucidating the SNR-t Bias of Diffusion Probabilistic Models

Meng Yu, Lei Sun, Jianhao Zeng +2

Diffusion Probabilistic Models have demonstrated remarkable performance across a wide range of generative tasks. However, we have observed that these models often suffer from a Sig…

cs.SD2021

Generalized Spatio-Temporal RNN Beamformer for Target Speech Separation

Yong Xu, Zhuohuang Zhang, Meng Yu +2

Although the conventional mask-based minimum variance distortionless response (MVDR) could reduce the non-linear distortion, the residual noise level of the MVDR separated speech i…

cs.SD2025

U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation

Xusheng Yang, Long Zhou, Wenfu Wang +6

We propose \textbf{U-Codec}, an \textbf{U}ltra low frame-rate neural speech \textbf{Codec} that achieves high-fidelity reconstruction and fast speech generation at an extremely low…

cs.CV2025

Frequency Regulation for Exposure Bias Mitigation in Diffusion Models

Meng Yu, Kun Zhan

Diffusion models exhibit impressive generative capabilities but are significantly impacted by exposure bias. In this paper, we make a key observation: the energy of predicted noisy…

cs.CL2026

LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

Hao Zhang, Weiwei Li, Rilin Chen +3

Achieving full-duplex communication in spoken dialogue systems (SDS) requires real-time coordination between listening, speaking, and thinking. This paper proposes a semantic voice…

eess.AS2021

Towards Robust Speaker Verification with Target Speaker Enhancement

Chunlei Zhang, Meng Yu, Chao Weng +1

This paper proposes the target speaker enhancement based speaker verification network (TASE-SVNet), an all neural model that couples target speaker enhancement and speaker embeddin…

cs.LG2019

A Unified Framework for Speech Separation

Fahimeh Bahmaninezhad, Shi-Xiong Zhang, Yong Xu +3

Speech separation refers to extracting each individual speech source in a given mixed signal. Recent advancements in speech separation and ongoing research in this area, have made…

eess.AS2023

Unifying Robustness and Fidelity: A Comprehensive Study of Pretrained Generative Methods for Speech Enhancement in Adverse Conditions

Heming Wang, Meng Yu, Hao Zhang +5

Enhancing speech signal quality in adverse acoustic environments is a persistent challenge in speech processing. Existing deep learning based enhancement methods often struggle to…

eess.AS2020

WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and Dereverberation

Zhaoheng Ni, Yong Xu, Meng Yu +4

This paper aims at eliminating the interfering speakers' speech, additive noise, and reverberation from the noisy multi-talker speech mixture that benefits automatic speech recogni…

eess.AS2023

Deep Audio Zooming: Beamwidth-Controllable Neural Beamformer

Meng Yu, Dong Yu

Audio zooming, a signal processing technique, enables selective focusing and enhancement of sound signals from a specified region, attenuating others. While traditional beamforming…

eess.AS2020

End-to-End Multi-Look Keyword Spotting

Meng Yu, Xuan Ji, Bo Wu +2

The performance of keyword spotting (KWS), measured in false alarms and false rejects, degrades significantly under the far field and noisy conditions. In this paper, we propose a…

eess.AS2022

NeuralEcho: A Self-Attentive Recurrent Neural Network For Unified Acoustic Echo Suppression And Speech Enhancement

Meng Yu, Yong Xu, Chunlei Zhang +2

Acoustic echo cancellation (AEC) plays an important role in the full-duplex speech communication as well as the front-end speech enhancement for recognition in the conditions when…

eess.AS2025

FNSE-SBGAN: Far-field Speech Enhancement with Schrodinger Bridge and Generative Adversarial Networks

Tong Lei, Qinwen Hu, Ziyao Lin +5

The prevailing method for neural speech enhancement predominantly utilizes fully-supervised deep learning with simulated pairs of far-field noisy-reverberant speech and clean speec…

cs.SD2026

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation

Tianxin Xie, Wentao Lei, Kai Jiang +27

Text-to-audio-video (T2AV) generation is central to applications such as filmmaking and world modeling. However, current models often fail to produce physically plausible sounds. P…

q-bio.MN2022

Identification of cancer-keeping genes as therapeutic targets by finding network control hubs

Xizhe Zhang, Chunyu Pan, Xinru Wei +9

Finding cancer driver genes has been a focal theme of cancer research and clinical studies. One of the recent approaches is based on network structural controllability that focuses…

eess.AS2022

EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers

Soumi Maiti, Yushi Ueda, Shinji Watanabe +4

In this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speak…

cs.SD2025

When Audio Generators Become Good Listeners: Generative Features for Understanding Tasks

Zeyu Xie, Chenxing Li, Xuenan Xu +6

This work pioneers the utilization of generative features in enhancing audio understanding. Unlike conventional discriminative features that directly optimize posterior and thus em…

astro-ph.IM2018

How shall we determine detection sensitivity in radio pulsar search?

Meng Yu

Determination of detection sensitivity in a number of previous pulsar search programmes was done via the straightfoward use of the radiometer equation. In the same surveys, the Fou…

eess.AS2023

Deep Neural Mel-Subband Beamformer for In-car Speech Separation

Vinay Kothapally, Yong Xu, Meng Yu +2

While current deep learning (DL)-based beamforming techniques have been proved effective in speech separation, they are often designed to process narrow-band (NB) frequencies indep…

cs.SD2026

Structure-Aware Piano Accompaniment via Style Planning and Dataset-Aligned Pattern Retrieval

Wanyu Zang, Yang Yu, Meng Yu

We introduce a structure-aware approach for symbolic piano accompaniment that decouples high-level planning from note-level realization. A lightweight transformer predicts an inter…

cs.CL2021

Joint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization

Brian Yan, Chunlei Zhang, Meng Yu +6

Conversational bilingual speech encompasses three types of utterances: two purely monolingual types and one intra-sententially code-switched type. In this work, we propose a genera…

cs.SD2026

Covo-Audio Technical Report

Wenfu Wang, Chenxing Li, Liqiang Zhang +23

In this work, we present Covo-Audio, a 7B-parameter end-to-end LALM that directly processes continuous audio inputs and generates audio outputs within a single unified architecture…

cs.CV2026

From Scale to Speed: Adaptive Test-Time Scaling for Image Editing

Xiangyan Qu, Zhenlong Yuan, Jing Tang +9

Image Chain-of-Thought (Image-CoT) is a test-time scaling paradigm that improves image generation by extending inference time. Most Image-CoT methods focus on text-to-image (T2I) g…

eess.AS2019

Time Domain Audio Visual Speech Separation

Jian Wu, Yong Xu, Shi-Xiong Zhang +4

Audio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new…

eess.AS2023

Hybrid AHS: A Hybrid of Kalman Filter and Deep Learning for Acoustic Howling Suppression

Hao Zhang, Meng Yu, Yuzhong Wu +2

Deep learning has been recently introduced for efficient acoustic howling suppression (AHS). However, the recurrent nature of howling creates a mismatch between offline training an…

cs.CR2017

TrustShadow: Secure Execution of Unmodified Applications with ARM TrustZone

Le Guan, Peng Liu, Xinyu Xing +4

The rapid evolution of Internet-of-Things (IoT) technologies has led to an emerging need to make it smarter. A variety of applications now run simultaneously on an ARM-based proces…

eess.AS2024

Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment

Yiwen Shao, Shi-Xiong Zhang, Yong Xu +4

In the field of multi-channel, multi-speaker Automatic Speech Recognition (ASR), the task of discerning and accurately transcribing a target speaker's speech within background nois…

cs.CV2024

VIFNet: An End-to-end Visible-Infrared Fusion Network for Image Dehazing

Meng Yu, Te Cui, Haoyang Lu +1

Image dehazing poses significant challenges in environmental perception. Recent research mainly focus on deep learning-based methods with single modality, while they may result in…

cs.SD2021

MIMO Self-attentive RNN Beamformer for Multi-speaker Speech Separation

Xiyun Li, Yong Xu, Meng Yu +4

Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance ove…

astro-ph.HE2018

Have we seen all glitches?

Meng Yu

Neutron star glitches are observed via artificially scheduled pulsar pulse arrival-time observations. Detection probability density of glitch events for a given data set is essenti…

eess.AS2023

Neural Network Augmented Kalman Filter for Robust Acoustic Howling Suppression

Yixuan Zhang, Hao Zhang, Meng Yu +1

Acoustic howling suppression (AHS) is a critical challenge in audio communication systems. In this paper, we propose a novel approach that leverages the power of neural networks (N…

eess.AS2021

ADL-MVDR: All deep learning MVDR beamformer for target speech separation

Zhuohuang Zhang, Yong Xu, Meng Yu +3

Speech separation algorithms are often used to separate the target speech from other interfering sources. However, purely neural network based speech separation systems often cause…

eess.AS2021

Deep Learning based Multi-Source Localization with Source Splitting and its Effectiveness in Multi-Talker Speech Recognition

Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe +2

Multi-source localization is an important and challenging technique for multi-talker conversation analysis. This paper proposes a novel supervised learning method using deep neural…

eess.AS2019

Improved Speaker-Dependent Separation for CHiME-5 Challenge

Jian Wu, Yong Xu, Shi-Xiong Zhang +4

This paper summarizes several follow-up contributions for improving our submitted NWPU speaker-dependent system for CHiME-5 challenge, which aims to solve the problem of multi-chan…

cs.CV2026

Bias mitigation in graph diffusion models

Meng Yu, Kun Zhan

Most existing graph diffusion models have significant bias problems. We observe that the forward diffusion's maximum perturbation distribution in most models deviates from the stan…

eess.AS2020

Enhancing End-to-End Multi-channel Speech Separation via Spatial Feature Learning

Rongzhi Gu, Shi-Xiong Zhang, Lianwu Chen +5

Hand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. Howe…

cs.SD2025

SMRU: Split-and-Merge Recurrent-based UNet for Acoustic Echo Cancellation and Noise Suppression

Zhihang Sun, Andong Li, Rilin Chen +4

The proliferation of deep neural networks has spawned the rapid development of acoustic echo cancellation and noise suppression, and plenty of prior arts have been proposed, which…

eess.AS2025

Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders

Lingling Dai, Andong Li, Tong Lei +3

Time-frequency (T-F) domain-based neural vocoders have shown promising results in synthesizing high-fidelity audio. Nevertheless, it remains unclear on the mechanism of effectively…

eess.AS2023

Deep Learning for Joint Acoustic Echo and Acoustic Howling Suppression in Hybrid Meetings

Hao Zhang, Meng Yu, Dong Yu

Hybrid meetings have become increasingly necessary during the post-COVID period and also brought new challenges for solving audio-related problems. In particular, the interplay bet…

cs.SD2019

A comprehensive study of speech separation: spectrogram vs waveform separation

Fahimeh Bahmaninezhad, Jian Wu, Rongzhi Gu +4

Speech separation has been studied widely for single-channel close-talk microphone recordings over the past few years; developed solutions are mostly in frequency-domain. Recently,…

cs.SD2025

From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models

Zhaoxi Mu, Rilin Chen, Andong Li +3

This paper introduces OmniGSE, a novel general speech enhancement (GSE) framework designed to mitigate the diverse distortions that speech signals encounter in real-world scenarios…

eess.AS2020

Distortionless Multi-Channel Target Speech Enhancement for Overlapped Speech Recognition

Bo Wu, Meng Yu, Lianwu Chen +4

Speech enhancement techniques based on deep learning have brought significant improvement on speech quality and intelligibility. Nevertheless, a large gain in speech quality measur…

eess.AS2020

Audio-Visual Speech Separation and Dereverberation with a Two-Stage Multimodal Network

Ke Tan, Yong Xu, Shi-Xiong Zhang +2

Background noise, interfering speech and room reverberation frequently distort target speech in real listening environments. In this study, we address joint speech separation and d…

eess.AS2019

Overlapped speech recognition from a jointly learned multi-channel neural speech extraction and representation

Bo Wu, Meng Yu, Lianwu Chen +3

We propose an end-to-end joint optimization framework of a multi-channel neural speech extraction and deep acoustic model without mel-filterbank (FBANK) extraction for overlapped s…

eess.AS2021

An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation

Daniel Michelsanti, Zheng-Hua Tan, Shi-Xiong Zhang +4

Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generat…

eess.AS2023

NeuralKalman: A Learnable Kalman Filter for Acoustic Echo Cancellation

Yixuan Zhang, Meng Yu, Hao Zhang +2

The robustness of the Kalman filter to double talk and its rapid convergence make it a popular approach for addressing acoustic echo cancellation (AEC) challenges. However, the ina…

eess.AS2025

SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

Helin Wang, Meng Yu, Jiarui Hai +5

In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot textbased speech editing and text-to-speech synthesis. S…

eess.AS2020

Neural Spatio-Temporal Beamformer for Target Speech Separation

Yong Xu, Meng Yu, Shi-Xiong Zhang +4

Purely neural network (NN) based speech separation and enhancement methods, although can achieve good objective scores, inevitably cause nonlinear speech distortions that are harmf…

astro-ph.HE2017

On the detection probability of neutron star glitches

Meng Yu, Qinjian Liu

Neutron stars are observed to undergo small, abrupt rotational speed-up. This phenomenon is known as glitch. In pulsar timing observations, detection of a neutron star glitch is co…

astro-ph.HE2009

Cooling and Heating Solid Quark Stars

Meng Yu

We present here a phenomenological solid quark star pulsar model to interpret the observed thermal X-ray emission of isolated pulsars. The heat capacity for solid quark stars was f…

cs.SD2026

DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer

Yisu Liu, Chenxing Li, Wanqian Zhang +6

Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and…

eess.AS2023

Advancing Acoustic Howling Suppression through Recursive Training of Neural Networks

Hao Zhang, Yixuan Zhang, Meng Yu +1

In this paper, we introduce a novel training framework designed to comprehensively address the acoustic howling issue by examining its fundamental formation process. This framework…

eess.AS2020

Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization

Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe +4

This paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end neural network manner, called directional automatic speech recognition (D-ASR), which…

cs.SD2018

Deep Extractor Network for Target Speaker Recovery From Single Channel Speech Mixtures

Jun Wang, Jie Chen, Dan Su +4

Speaker-aware source separation methods are promising workarounds for major difficulties such as arbitrary source permutation and unknown number of sources. However, it remains cha…

eess.AS2024

Restorative Speech Enhancement: A Progressive Approach Using SE and Codec Modules

Hsin-Tien Chiang, Hao Zhang, Yong Xu +2

In challenging environments with significant noise and reverberation, traditional speech enhancement (SE) methods often lead to over-suppressed speech, creating artifacts during li…

cs.SD2021

Improving RNN Transducer With Target Speaker Extraction and Neural Uncertainty Estimation

Jiatong Shi, Chunlei Zhang, Chao Weng +3

Target-speaker speech recognition aims to recognize target-speaker speech from noisy environments with background noise and interfering speakers. This work presents a joint framewo…

eess.AS2024

Neural Ambisonic Encoding For Multi-Speaker Scenarios Using A Circular Microphone Array

Yue Qiao, Vinay Kothapally, Meng Yu +1

Spatial audio formats like Ambisonics are playback device layout-agnostic and well-suited for applications such as teleconferencing and virtual reality. Conventional Ambisonic enco…

math.AP2024

Sharp extinction rates for positive solutions of fast diffusion equations

Tobias König, Meng Yu

Let and . It is known that positive solutions to the (fractional) fast diffusion equation on $(0, \infty)…

cs.SD2022

Enhancing Zero-Shot Many to Many Voice Conversion with Self-Attention VAE

Ziang Long, Yunling Zheng, Meng Yu +1

Variational auto-encoder (VAE) is an effective neural network architecture to disentangle a speech utterance into speaker identity and linguistic content latent embeddings, then ge…

cs.CV2024

Open-RGBT: Open-vocabulary RGB-T Zero-shot Semantic Segmentation in Open-world Environments

Meng Yu, Luojie Yang, Xunjie He +2

Semantic segmentation is a critical technique for effective scene understanding. Traditional RGB-T semantic segmentation models often struggle to generalize across diverse scenario…

cs.SD2025

Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning

Shu Wu, Chenxing Li, Wenfu Wang +4

Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities thr…

cs.CL2019

DurIAN: Duration Informed Attention Network For Multimodal Synthesis

Chengzhu Yu, Heng Lu, Na Hu +9

In this paper, we present a generic and robust multimodal synthesis system that produces highly natural speech and facial expression simultaneously. The key component of this syste…

cs.CV2026

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

Yingmao Miao, Pengfei Zhang, Xiaochen Lv +5

While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuri…

eess.AS2022

Joint Neural AEC and Beamforming with Double-Talk Detection

Vinay Kothapally, Yong Xu, Meng Yu +2

Acoustic echo cancellation (AEC) in full-duplex communication systems eliminates acoustic feedback. However, nonlinear distortions induced by audio devices, background noise, rever…

cs.SD2022

FAST-RIR: Fast neural diffuse room impulse response generator

Anton Ratnarajah, Shi-Xiong Zhang, Meng Yu +3

We present a neural-network-based fast diffuse room impulse response generator (FAST-RIR) for generating room impulse responses (RIRs) for a given acoustic environment. Our FAST-RI…

eess.AS2021

MetricNet: Towards Improved Modeling For Non-Intrusive Speech Quality Assessment

Meng Yu, Chunlei Zhang, Yong Xu +2

The objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech q…

astro-ph.HE2019

Tracking the footprints of the radio pulsar B172747: proper motion, host supernova remnant, and the glitches

Peter Shternin, Aida Kirichenko, Dmitry Zyuzin +4

The bright radio pulsar B172747 with a characteristic age of 80 kyr is among the first pulsars discovered 50 yr ago. Using its regular timing observations and interferometric po…

cs.SD2019

End-to-End Multi-Channel Speech Separation

Rongzhi Gu, Jian Wu, Shi-Xiong Zhang +6

The end-to-end approach for single-channel speech separation has been studied recently and shown promising results. This paper extended the previous approach and proposed a new end…

astro-ph.HE2019

Study of Three Rotating Radio Transients with FAST

Jiguang Lu, Bo Peng, Kuo Liu +7

Rotating radio transients (RRATs) are peculiar astronomical objects whose emission mechanism remains under investigation. In this paper, we present observations of three RRATs, J15…

astro-ph.HE2019

FAST ultra-wideband observation of abnormal emission-shift events of PSR B0919+06

Ye-Zhao Yu, Bo Peng, Kuo Liu +6

PSR B0919+06 is known for its abnormal emission phenomenon, where the pulse emission window occasionally shifts progressively in longitude and returns afterwards. The physical mech…

astro-ph.HE2012

Glitches detected in southern radio pulsars

Meng Yu

Parkes pulse arrival-time data for 165 radio pulsars spanning from 1990 to 2011 have been searched for period glitches. Forty-six events out of the detected 107 glitches were found…

cs.CV2025

N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models

Yuxin Wang, Lei Ke, Boqiang Zhang +6

While current multimodal models can answer questions based on 2D images, they lack intrinsic 3D object perception, limiting their ability to comprehend spatial relationships and de…

econ.TH2019

Equilibrium in Production Chains with Multiple Upstream Partners

Meng Yu, Junnan Zhang

In this paper, we extend and improve the production chain model introduced by Kikuchi et al. (2018). Utilizing the theory of monotone concave operators, we prove the existence, uni…

cs.CV2024

Joint Conditional Diffusion Model for Image Restoration with Mixed Degradations

Yufeng Yue, Meng Yu, Luojie Yang +1

Image restoration is rather challenging in adverse weather conditions, especially when multiple degradations occur simultaneously. Blind image decomposition was proposed to tackle…

eess.AS2023

Deep AHS: A Deep Learning Approach to Acoustic Howling Suppression

Hao Zhang, Meng Yu, Dong Yu

In this paper, we formulate acoustic howling suppression (AHS) as a supervised learning problem and propose a deep learning approach, called Deep AHS, to address it. Deep AHS is tr…

eess.AS2021

Self-supervised Text-independent Speaker Verification using Prototypical Momentum Contrastive Learning

Wei Xia, Chunlei Zhang, Chao Weng +2

In this study, we investigate self-supervised representation learning for speaker verification (SV). First, we examine a simple contrastive learning approach (SimCLR) with a moment…

astro-ph.HE2019

The Radiation Structure of PSR B201628 Observed with FAST

Jiguang Lu, Bo Peng, Renxin Xu +8

With the largest dish Five-hundred-meter Aperture Spherical radio Telescope (FAST), both the mean and single pulses of PSR B201628, especially including the single-pulse structu…

cs.CV2025

TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation Models

Meng Yu, Te Cui, Qitong Chu +3

Reliable semantic segmentation of open environments is essential for intelligent systems, yet significant problems remain: 1) Existing RGB-T semantic segmentation models mainly rel…

eess.AS2021

TeCANet: Temporal-Contextual Attention Network for Environment-Aware Speech Dereverberation

Helin Wang, Bo Wu, Lianwu Chen +7

In this paper, we exploit the effective way to leverage contextual information to improve the speech dereverberation performance in real-world reverberant environments. We propose…

eess.AS2021

Multi-channel Multi-frame ADL-MVDR for Target Speech Separation

Zhuohuang Zhang, Yong Xu, Meng Yu +4

Many purely neural network based speech separation approaches have been proposed to improve objective assessment scores, but they often introduce nonlinear distortions that are har…

cs.SD2026

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

Jiliang Hu, Wenfu Wang, Zuchao Li +6

Recent advances in large audio language models (LALMs) have greatly enhanced multimodal conversational systems. However, existing benchmarks remain limited -- they are mainly Engli…

cs.CL2025

AzeroS: Extending LLM to Speech with Self-Generated Instruction-Free Tuning

Yiwen Shao, Wei Liu, Jiahong Li +4

Extending large language models (LLMs) to the speech domain has recently gained significant attention. A typical approach connects a pretrained LLM with an audio encoder through a…

cs.SD2025

BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective

Andong Li, Tong Lei, Rilin Chen +5

This paper revisits the neural vocoder task through the lens of audio restoration and propose a novel diffusion vocoder called BridgeVoC. Specifically, by rank analysis, we compare…

cs.SD2025

Target matching based generative model for speech enhancement

Taihui Wang, Rilin Chen, Tong Lei +4

The design of mean and variance schedules for the perturbed signal is a fundamental challenge in generative models. While score-based and Schrödinger bridge-based models require c…