3 papers
cs.SD2026
Audio Spotforming via Post-Filtering Using Cross-Array Non-target Estimates
Yuto Ishikawa, Li Li, Shogo Seki +1
Audio spotforming is a technique for extracting target speech from noisy mixtures by utilizing multiple microphone arrays. Conventional methods estimate a shared target speech comp…
eess.AS2026
DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models
Li Li, Ming Cheng, Weixin Zhu +3
Multi-speaker automatic speech recognition (ASR) aims to transcribe conversational speech involving multiple speakers, requiring the model to capture not only what was said, but al…
cs.SD2025
Improving DF-Conformer Using Hydra For High-Fidelity Generative Speech Enhancement on Discrete Codec Token
Shogo Seki, Shaoxiang Dang, Li Li
The Dilated FAVOR Conformer (DF-Conformer) is an efficient variant of the Conformer architecture designed for speech enhancement (SE). It employs fast attention through positive or…