activity
20242026
most citedExploration of Adapter for Noise Robust Automatic Speech Recognition

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2026

Grounded Decoding for Autoregressive Speech Enhancement via Adaptive Code-Space Grounding and Local LLM Refinement

Hao Shi, Yuan Gao, Zhaoheng Ni +4

Large language model (LLM)-based autoregressive speech enhancement (SE) produces natural speech using learned clean-speech priors, but may hallucinate content unsupported by the in…

cs.SD2026

Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition

Hao Shi, Yuan Gao, Xugang Lu +1

Large Language Models (LLMs) are effective decoders for Serialized Output Training (SOT) in two-talker automatic speech recognition (ASR), but their performance degrades substantia…

cs.SD2025

Combining Deterministic Enhanced Conditions with Dual-Streaming Encoding for Diffusion-Based Speech Enhancement

Hao Shi, Xugang Lu, Kazuki Shimada +1

Diffusion-based speech enhancement (SE) models need to incorporate correct prior knowledge as reliable conditions to generate accurate predictions. However, providing reliable cond…

cs.SD2025

Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network

Yuan Gao, Hao Shi, Yahui Fu +2

This study investigates the interaction between personality traits and emotion expression, exploring how personality information can improve speech emotion recognition (SER). We co…

cs.SD2024

Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition

Hao Shi, Yuan Gao, Zhaoheng Ni +1

Serialized output training (SOT) attracts increasing attention due to its convenience and flexibility for multi-speaker automatic speech recognition (ASR). However, it is not easy…

cs.SD20241 cited

Exploration of Adapter for Noise Robust Automatic Speech Recognition

Hao Shi, Tatsuya Kawahara

Adapting an automatic speech recognition (ASR) system to unseen noise environments is crucial. Integrating adapters into neural networks has emerged as a potent technique for trans…