1 citations · 1 across the 7 of their papers we have counts for
6 papers · 1 filter
Grounded Decoding for Autoregressive Speech Enhancement via Adaptive Code-Space Grounding and Local LLM Refinement
Hao Shi, Yuan Gao, Zhaoheng Ni +4
Large language model (LLM)-based autoregressive speech enhancement (SE) produces natural speech using learned clean-speech priors, but may hallucinate content unsupported by the in…
Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition
Hao Shi, Yuan Gao, Xugang Lu +1
Large Language Models (LLMs) are effective decoders for Serialized Output Training (SOT) in two-talker automatic speech recognition (ASR), but their performance degrades substantia…
Combining Deterministic Enhanced Conditions with Dual-Streaming Encoding for Diffusion-Based Speech Enhancement
Hao Shi, Xugang Lu, Kazuki Shimada +1
Diffusion-based speech enhancement (SE) models need to incorporate correct prior knowledge as reliable conditions to generate accurate predictions. However, providing reliable cond…
Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network
Yuan Gao, Hao Shi, Yahui Fu +2
This study investigates the interaction between personality traits and emotion expression, exploring how personality information can improve speech emotion recognition (SER). We co…
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
Hao Shi, Yuan Gao, Zhaoheng Ni +1
Serialized output training (SOT) attracts increasing attention due to its convenience and flexibility for multi-speaker automatic speech recognition (ASR). However, it is not easy…
Exploration of Adapter for Noise Robust Automatic Speech Recognition
Hao Shi, Tatsuya Kawahara
Adapting an automatic speech recognition (ASR) system to unseen noise environments is crucial. Integrating adapters into neural networks has emerged as a potent technique for trans…