activity
20242026
collaborators

5 papers

cs.SD2026

Audio Spotforming via Post-Filtering Using Cross-Array Non-target Estimates

Yuto Ishikawa, Li Li, Shogo Seki +1

Audio spotforming is a technique for extracting target speech from noisy mixtures by utilizing multiple microphone arrays. Conventional methods estimate a shared target speech comp…

eess.AS2026

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models

Li Li, Ming Cheng, Weixin Zhu +3

Multi-speaker automatic speech recognition (ASR) aims to transcribe conversational speech involving multiple speakers, requiring the model to capture not only what was said, but al…

cs.SD2025

Improving DF-Conformer Using Hydra For High-Fidelity Generative Speech Enhancement on Discrete Codec Token

Shogo Seki, Shaoxiang Dang, Li Li

The Dilated FAVOR Conformer (DF-Conformer) is an efficient variant of the Conformer architecture designed for speech enhancement (SE). It employs fast attention through positive or…

cs.SD2024

Audio Spotforming Using Nonnegative Tensor Factorization with Attractor-Based Regularization

Shoma Ayano, Li Li, Shogo Seki +1

Spotforming is a target-speaker extraction technique that uses multiple microphone arrays. This method applies beamforming (BF) to each microphone array, and the common components…

cs.SD2024

Improved Remixing Process for Domain Adaptation-Based Speech Enhancement by Mitigating Data Imbalance in Signal-to-Noise Ratio

Li Li, Shogo Seki

RemixIT and Remixed2Remixed are domain adaptation-based speech enhancement (DASE) methods that use a teacher model trained in full supervision to generate pseudo-paired data by rem…