Publications (85)
Harmonic Summing Improves Pulsar Detection Sensitivity: A Probability Analysis
Meng Yu
Practical application of the harmonic summing technique in the power-spectrum analysis for searching pulsars has exhibited the technique's effectiveness. In this paper, theoretical…
Elucidating the SNR-t Bias of Diffusion Probabilistic Models
Meng Yu, Lei Sun, Jianhao Zeng +2
Diffusion Probabilistic Models have demonstrated remarkable performance across a wide range of generative tasks. However, we have observed that these models often suffer from a Sig…
Generalized Spatio-Temporal RNN Beamformer for Target Speech Separation
Yong Xu, Zhuohuang Zhang, Meng Yu +2
Although the conventional mask-based minimum variance distortionless response (MVDR) could reduce the non-linear distortion, the residual noise level of the MVDR separated speech i…
U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation
Xusheng Yang, Long Zhou, Wenfu Wang +6
We propose \textbf{U-Codec}, an \textbf{U}ltra low frame-rate neural speech \textbf{Codec} that achieves high-fidelity reconstruction and fast speech generation at an extremely low…
Frequency Regulation for Exposure Bias Mitigation in Diffusion Models
Meng Yu, Kun Zhan
Diffusion models exhibit impressive generative capabilities but are significantly impacted by exposure bias. In this paper, we make a key observation: the energy of predicted noisy…
LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
Hao Zhang, Weiwei Li, Rilin Chen +3
Achieving full-duplex communication in spoken dialogue systems (SDS) requires real-time coordination between listening, speaking, and thinking. This paper proposes a semantic voice…
Towards Robust Speaker Verification with Target Speaker Enhancement
Chunlei Zhang, Meng Yu, Chao Weng +1
This paper proposes the target speaker enhancement based speaker verification network (TASE-SVNet), an all neural model that couples target speaker enhancement and speaker embeddin…
A Unified Framework for Speech Separation
Fahimeh Bahmaninezhad, Shi-Xiong Zhang, Yong Xu +3
Speech separation refers to extracting each individual speech source in a given mixed signal. Recent advancements in speech separation and ongoing research in this area, have made…
Unifying Robustness and Fidelity: A Comprehensive Study of Pretrained Generative Methods for Speech Enhancement in Adverse Conditions
Heming Wang, Meng Yu, Hao Zhang +5
Enhancing speech signal quality in adverse acoustic environments is a persistent challenge in speech processing. Existing deep learning based enhancement methods often struggle to…
WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and Dereverberation
Zhaoheng Ni, Yong Xu, Meng Yu +4
This paper aims at eliminating the interfering speakers' speech, additive noise, and reverberation from the noisy multi-talker speech mixture that benefits automatic speech recogni…
Deep Audio Zooming: Beamwidth-Controllable Neural Beamformer
Meng Yu, Dong Yu
Audio zooming, a signal processing technique, enables selective focusing and enhancement of sound signals from a specified region, attenuating others. While traditional beamforming…
End-to-End Multi-Look Keyword Spotting
Meng Yu, Xuan Ji, Bo Wu +2
The performance of keyword spotting (KWS), measured in false alarms and false rejects, degrades significantly under the far field and noisy conditions. In this paper, we propose a…
NeuralEcho: A Self-Attentive Recurrent Neural Network For Unified Acoustic Echo Suppression And Speech Enhancement
Meng Yu, Yong Xu, Chunlei Zhang +2
Acoustic echo cancellation (AEC) plays an important role in the full-duplex speech communication as well as the front-end speech enhancement for recognition in the conditions when…
FNSE-SBGAN: Far-field Speech Enhancement with Schrodinger Bridge and Generative Adversarial Networks
Tong Lei, Qinwen Hu, Ziyao Lin +5
The prevailing method for neural speech enhancement predominantly utilizes fully-supervised deep learning with simulated pairs of far-field noisy-reverberant speech and clean speec…
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation
Tianxin Xie, Wentao Lei, Kai Jiang +27
Text-to-audio-video (T2AV) generation is central to applications such as filmmaking and world modeling. However, current models often fail to produce physically plausible sounds. P…
Identification of cancer-keeping genes as therapeutic targets by finding network control hubs
Xizhe Zhang, Chunyu Pan, Xinru Wei +9
Finding cancer driver genes has been a focal theme of cancer research and clinical studies. One of the recent approaches is based on network structural controllability that focuses…
EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers
Soumi Maiti, Yushi Ueda, Shinji Watanabe +4
In this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speak…
When Audio Generators Become Good Listeners: Generative Features for Understanding Tasks
Zeyu Xie, Chenxing Li, Xuenan Xu +6
This work pioneers the utilization of generative features in enhancing audio understanding. Unlike conventional discriminative features that directly optimize posterior and thus em…
How shall we determine detection sensitivity in radio pulsar search?
Meng Yu
Determination of detection sensitivity in a number of previous pulsar search programmes was done via the straightfoward use of the radiometer equation. In the same surveys, the Fou…
Deep Neural Mel-Subband Beamformer for In-car Speech Separation
Vinay Kothapally, Yong Xu, Meng Yu +2
While current deep learning (DL)-based beamforming techniques have been proved effective in speech separation, they are often designed to process narrow-band (NB) frequencies indep…
Structure-Aware Piano Accompaniment via Style Planning and Dataset-Aligned Pattern Retrieval
Wanyu Zang, Yang Yu, Meng Yu
We introduce a structure-aware approach for symbolic piano accompaniment that decouples high-level planning from note-level realization. A lightweight transformer predicts an inter…
Joint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization
Brian Yan, Chunlei Zhang, Meng Yu +6
Conversational bilingual speech encompasses three types of utterances: two purely monolingual types and one intra-sententially code-switched type. In this work, we propose a genera…
Covo-Audio Technical Report
Wenfu Wang, Chenxing Li, Liqiang Zhang +23
In this work, we present Covo-Audio, a 7B-parameter end-to-end LALM that directly processes continuous audio inputs and generates audio outputs within a single unified architecture…
From Scale to Speed: Adaptive Test-Time Scaling for Image Editing
Xiangyan Qu, Zhenlong Yuan, Jing Tang +9
Image Chain-of-Thought (Image-CoT) is a test-time scaling paradigm that improves image generation by extending inference time. Most Image-CoT methods focus on text-to-image (T2I) g…
Time Domain Audio Visual Speech Separation
Jian Wu, Yong Xu, Shi-Xiong Zhang +4
Audio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new…
Hybrid AHS: A Hybrid of Kalman Filter and Deep Learning for Acoustic Howling Suppression
Hao Zhang, Meng Yu, Yuzhong Wu +2
Deep learning has been recently introduced for efficient acoustic howling suppression (AHS). However, the recurrent nature of howling creates a mismatch between offline training an…
TrustShadow: Secure Execution of Unmodified Applications with ARM TrustZone
Le Guan, Peng Liu, Xinyu Xing +4
The rapid evolution of Internet-of-Things (IoT) technologies has led to an emerging need to make it smarter. A variety of applications now run simultaneously on an ARM-based proces…
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
Yiwen Shao, Shi-Xiong Zhang, Yong Xu +4
In the field of multi-channel, multi-speaker Automatic Speech Recognition (ASR), the task of discerning and accurately transcribing a target speaker's speech within background nois…
VIFNet: An End-to-end Visible-Infrared Fusion Network for Image Dehazing
Meng Yu, Te Cui, Haoyang Lu +1
Image dehazing poses significant challenges in environmental perception. Recent research mainly focus on deep learning-based methods with single modality, while they may result in…
MIMO Self-attentive RNN Beamformer for Multi-speaker Speech Separation
Xiyun Li, Yong Xu, Meng Yu +4
Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance ove…
Have we seen all glitches?
Meng Yu
Neutron star glitches are observed via artificially scheduled pulsar pulse arrival-time observations. Detection probability density of glitch events for a given data set is essenti…
Neural Network Augmented Kalman Filter for Robust Acoustic Howling Suppression
Yixuan Zhang, Hao Zhang, Meng Yu +1
Acoustic howling suppression (AHS) is a critical challenge in audio communication systems. In this paper, we propose a novel approach that leverages the power of neural networks (N…
ADL-MVDR: All deep learning MVDR beamformer for target speech separation
Zhuohuang Zhang, Yong Xu, Meng Yu +3
Speech separation algorithms are often used to separate the target speech from other interfering sources. However, purely neural network based speech separation systems often cause…
Deep Learning based Multi-Source Localization with Source Splitting and its Effectiveness in Multi-Talker Speech Recognition
Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe +2
Multi-source localization is an important and challenging technique for multi-talker conversation analysis. This paper proposes a novel supervised learning method using deep neural…
Improved Speaker-Dependent Separation for CHiME-5 Challenge
Jian Wu, Yong Xu, Shi-Xiong Zhang +4
This paper summarizes several follow-up contributions for improving our submitted NWPU speaker-dependent system for CHiME-5 challenge, which aims to solve the problem of multi-chan…
Bias mitigation in graph diffusion models
Meng Yu, Kun Zhan
Most existing graph diffusion models have significant bias problems. We observe that the forward diffusion's maximum perturbation distribution in most models deviates from the stan…
Enhancing End-to-End Multi-channel Speech Separation via Spatial Feature Learning
Rongzhi Gu, Shi-Xiong Zhang, Lianwu Chen +5
Hand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. Howe…
SMRU: Split-and-Merge Recurrent-based UNet for Acoustic Echo Cancellation and Noise Suppression
Zhihang Sun, Andong Li, Rilin Chen +4
The proliferation of deep neural networks has spawned the rapid development of acoustic echo cancellation and noise suppression, and plenty of prior arts have been proposed, which…
Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders
Lingling Dai, Andong Li, Tong Lei +3
Time-frequency (T-F) domain-based neural vocoders have shown promising results in synthesizing high-fidelity audio. Nevertheless, it remains unclear on the mechanism of effectively…
Deep Learning for Joint Acoustic Echo and Acoustic Howling Suppression in Hybrid Meetings
Hao Zhang, Meng Yu, Dong Yu
Hybrid meetings have become increasingly necessary during the post-COVID period and also brought new challenges for solving audio-related problems. In particular, the interplay bet…
A comprehensive study of speech separation: spectrogram vs waveform separation
Fahimeh Bahmaninezhad, Jian Wu, Rongzhi Gu +4
Speech separation has been studied widely for single-channel close-talk microphone recordings over the past few years; developed solutions are mostly in frequency-domain. Recently,…
From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
Zhaoxi Mu, Rilin Chen, Andong Li +3
This paper introduces OmniGSE, a novel general speech enhancement (GSE) framework designed to mitigate the diverse distortions that speech signals encounter in real-world scenarios…
Distortionless Multi-Channel Target Speech Enhancement for Overlapped Speech Recognition
Bo Wu, Meng Yu, Lianwu Chen +4
Speech enhancement techniques based on deep learning have brought significant improvement on speech quality and intelligibility. Nevertheless, a large gain in speech quality measur…
Audio-Visual Speech Separation and Dereverberation with a Two-Stage Multimodal Network
Ke Tan, Yong Xu, Shi-Xiong Zhang +2
Background noise, interfering speech and room reverberation frequently distort target speech in real listening environments. In this study, we address joint speech separation and d…
Overlapped speech recognition from a jointly learned multi-channel neural speech extraction and representation
Bo Wu, Meng Yu, Lianwu Chen +3
We propose an end-to-end joint optimization framework of a multi-channel neural speech extraction and deep acoustic model without mel-filterbank (FBANK) extraction for overlapped s…
An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation
Daniel Michelsanti, Zheng-Hua Tan, Shi-Xiong Zhang +4
Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generat…
NeuralKalman: A Learnable Kalman Filter for Acoustic Echo Cancellation
Yixuan Zhang, Meng Yu, Hao Zhang +2
The robustness of the Kalman filter to double talk and its rapid convergence make it a popular approach for addressing acoustic echo cancellation (AEC) challenges. However, the ina…
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
Helin Wang, Meng Yu, Jiarui Hai +5
In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot textbased speech editing and text-to-speech synthesis. S…
Neural Spatio-Temporal Beamformer for Target Speech Separation
Yong Xu, Meng Yu, Shi-Xiong Zhang +4
Purely neural network (NN) based speech separation and enhancement methods, although can achieve good objective scores, inevitably cause nonlinear speech distortions that are harmf…
On the detection probability of neutron star glitches
Meng Yu, Qinjian Liu
Neutron stars are observed to undergo small, abrupt rotational speed-up. This phenomenon is known as glitch. In pulsar timing observations, detection of a neutron star glitch is co…
Cooling and Heating Solid Quark Stars
Meng Yu
We present here a phenomenological solid quark star pulsar model to interpret the observed thermal X-ray emission of isolated pulsars. The heat capacity for solid quark stars was f…
DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer
Yisu Liu, Chenxing Li, Wanqian Zhang +6
Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and…
Advancing Acoustic Howling Suppression through Recursive Training of Neural Networks
Hao Zhang, Yixuan Zhang, Meng Yu +1
In this paper, we introduce a novel training framework designed to comprehensively address the acoustic howling issue by examining its fundamental formation process. This framework…
Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization
Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe +4
This paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end neural network manner, called directional automatic speech recognition (D-ASR), which…
Deep Extractor Network for Target Speaker Recovery From Single Channel Speech Mixtures
Jun Wang, Jie Chen, Dan Su +4
Speaker-aware source separation methods are promising workarounds for major difficulties such as arbitrary source permutation and unknown number of sources. However, it remains cha…
Restorative Speech Enhancement: A Progressive Approach Using SE and Codec Modules
Hsin-Tien Chiang, Hao Zhang, Yong Xu +2
In challenging environments with significant noise and reverberation, traditional speech enhancement (SE) methods often lead to over-suppressed speech, creating artifacts during li…
Improving RNN Transducer With Target Speaker Extraction and Neural Uncertainty Estimation
Jiatong Shi, Chunlei Zhang, Chao Weng +3
Target-speaker speech recognition aims to recognize target-speaker speech from noisy environments with background noise and interfering speakers. This work presents a joint framewo…
Neural Ambisonic Encoding For Multi-Speaker Scenarios Using A Circular Microphone Array
Yue Qiao, Vinay Kothapally, Meng Yu +1
Spatial audio formats like Ambisonics are playback device layout-agnostic and well-suited for applications such as teleconferencing and virtual reality. Conventional Ambisonic enco…
Sharp extinction rates for positive solutions of fast diffusion equations
Tobias König, Meng Yu
Let and . It is known that positive solutions to the (fractional) fast diffusion equation on $(0, \infty)…
Enhancing Zero-Shot Many to Many Voice Conversion with Self-Attention VAE
Ziang Long, Yunling Zheng, Meng Yu +1
Variational auto-encoder (VAE) is an effective neural network architecture to disentangle a speech utterance into speaker identity and linguistic content latent embeddings, then ge…
Open-RGBT: Open-vocabulary RGB-T Zero-shot Semantic Segmentation in Open-world Environments
Meng Yu, Luojie Yang, Xunjie He +2
Semantic segmentation is a critical technique for effective scene understanding. Traditional RGB-T semantic segmentation models often struggle to generalize across diverse scenario…
Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning
Shu Wu, Chenxing Li, Wenfu Wang +4
Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities thr…
DurIAN: Duration Informed Attention Network For Multimodal Synthesis
Chengzhu Yu, Heng Lu, Na Hu +9
In this paper, we present a generic and robust multimodal synthesis system that produces highly natural speech and facial expression simultaneously. The key component of this syste…
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
Yingmao Miao, Pengfei Zhang, Xiaochen Lv +5
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuri…
Joint Neural AEC and Beamforming with Double-Talk Detection
Vinay Kothapally, Yong Xu, Meng Yu +2
Acoustic echo cancellation (AEC) in full-duplex communication systems eliminates acoustic feedback. However, nonlinear distortions induced by audio devices, background noise, rever…
FAST-RIR: Fast neural diffuse room impulse response generator
Anton Ratnarajah, Shi-Xiong Zhang, Meng Yu +3
We present a neural-network-based fast diffuse room impulse response generator (FAST-RIR) for generating room impulse responses (RIRs) for a given acoustic environment. Our FAST-RI…
MetricNet: Towards Improved Modeling For Non-Intrusive Speech Quality Assessment
Meng Yu, Chunlei Zhang, Yong Xu +2
The objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech q…
Tracking the footprints of the radio pulsar B172747: proper motion, host supernova remnant, and the glitches
Peter Shternin, Aida Kirichenko, Dmitry Zyuzin +4
The bright radio pulsar B172747 with a characteristic age of 80 kyr is among the first pulsars discovered 50 yr ago. Using its regular timing observations and interferometric po…
End-to-End Multi-Channel Speech Separation
Rongzhi Gu, Jian Wu, Shi-Xiong Zhang +6
The end-to-end approach for single-channel speech separation has been studied recently and shown promising results. This paper extended the previous approach and proposed a new end…
Study of Three Rotating Radio Transients with FAST
Jiguang Lu, Bo Peng, Kuo Liu +7
Rotating radio transients (RRATs) are peculiar astronomical objects whose emission mechanism remains under investigation. In this paper, we present observations of three RRATs, J15…
FAST ultra-wideband observation of abnormal emission-shift events of PSR B0919+06
Ye-Zhao Yu, Bo Peng, Kuo Liu +6
PSR B0919+06 is known for its abnormal emission phenomenon, where the pulse emission window occasionally shifts progressively in longitude and returns afterwards. The physical mech…
Glitches detected in southern radio pulsars
Meng Yu
Parkes pulse arrival-time data for 165 radio pulsars spanning from 1990 to 2011 have been searched for period glitches. Forty-six events out of the detected 107 glitches were found…
N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
Yuxin Wang, Lei Ke, Boqiang Zhang +6
While current multimodal models can answer questions based on 2D images, they lack intrinsic 3D object perception, limiting their ability to comprehend spatial relationships and de…
Equilibrium in Production Chains with Multiple Upstream Partners
Meng Yu, Junnan Zhang
In this paper, we extend and improve the production chain model introduced by Kikuchi et al. (2018). Utilizing the theory of monotone concave operators, we prove the existence, uni…
Joint Conditional Diffusion Model for Image Restoration with Mixed Degradations
Yufeng Yue, Meng Yu, Luojie Yang +1
Image restoration is rather challenging in adverse weather conditions, especially when multiple degradations occur simultaneously. Blind image decomposition was proposed to tackle…
Deep AHS: A Deep Learning Approach to Acoustic Howling Suppression
Hao Zhang, Meng Yu, Dong Yu
In this paper, we formulate acoustic howling suppression (AHS) as a supervised learning problem and propose a deep learning approach, called Deep AHS, to address it. Deep AHS is tr…
Self-supervised Text-independent Speaker Verification using Prototypical Momentum Contrastive Learning
Wei Xia, Chunlei Zhang, Chao Weng +2
In this study, we investigate self-supervised representation learning for speaker verification (SV). First, we examine a simple contrastive learning approach (SimCLR) with a moment…
The Radiation Structure of PSR B201628 Observed with FAST
Jiguang Lu, Bo Peng, Renxin Xu +8
With the largest dish Five-hundred-meter Aperture Spherical radio Telescope (FAST), both the mean and single pulses of PSR B201628, especially including the single-pulse structu…
TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation Models
Meng Yu, Te Cui, Qitong Chu +3
Reliable semantic segmentation of open environments is essential for intelligent systems, yet significant problems remain: 1) Existing RGB-T semantic segmentation models mainly rel…
TeCANet: Temporal-Contextual Attention Network for Environment-Aware Speech Dereverberation
Helin Wang, Bo Wu, Lianwu Chen +7
In this paper, we exploit the effective way to leverage contextual information to improve the speech dereverberation performance in real-world reverberant environments. We propose…
Multi-channel Multi-frame ADL-MVDR for Target Speech Separation
Zhuohuang Zhang, Yong Xu, Meng Yu +4
Many purely neural network based speech separation approaches have been proposed to improve objective assessment scores, but they often introduce nonlinear distortions that are har…
VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents
Jiliang Hu, Wenfu Wang, Zuchao Li +6
Recent advances in large audio language models (LALMs) have greatly enhanced multimodal conversational systems. However, existing benchmarks remain limited -- they are mainly Engli…
AzeroS: Extending LLM to Speech with Self-Generated Instruction-Free Tuning
Yiwen Shao, Wei Liu, Jiahong Li +4
Extending large language models (LLMs) to the speech domain has recently gained significant attention. A typical approach connects a pretrained LLM with an audio encoder through a…
BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective
Andong Li, Tong Lei, Rilin Chen +5
This paper revisits the neural vocoder task through the lens of audio restoration and propose a novel diffusion vocoder called BridgeVoC. Specifically, by rank analysis, we compare…
Target matching based generative model for speech enhancement
Taihui Wang, Rilin Chen, Tong Lei +4
The design of mean and variance schedules for the perturbed signal is a fundamental challenge in generative models. While score-based and Schrödinger bridge-based models require c…