activity
20212026
collaborators
Showing cs.SDShow all

7 papers · 1 filter

cs.SD2026

Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization

Shuhai Peng, Jinjiang Liu, Hui Lu +5

Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends…

cs.SD2026

StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model

Shuhai Peng, Hui Lu, Jinjiang Liu +8

While generative models have set new benchmarks for Target Speaker Extraction (TSE), their inherent reliance on global context precludes deployment in real-time applications. Direc…

cs.SD2025

LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation

Jun Chen, Shichao Hu, Jiuxin Lin +8

In-car multi-zone speech separation, which captures voices from different speech zones, plays a crucial role in human-vehicle interaction. Although previous SpatialNet has achieved…

cs.SD2025

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding

Zijian Lin, Yang Zhang, Yougen Yuan +5

Modern autoregressive speech synthesis models leveraging language models have demonstrated remarkable performance. However, the sequential nature of next token prediction in these…

cs.SD2024

3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications

Shulin He, Jinjiang liu, Hao Li +3

Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current…

cs.SD2022

Speech Enhancement with Intelligent Neural Homomorphic Synthesis

Shulin He, Wei Rao, Jinjiang Liu +5

Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a…