From the 1 of 4 linked papers with an AI index.
4 papers
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
Chih-Ning Chen, Jen-Cheng Hou, Hsin-Min Wang +3
The paper introduces a reinforcement learning framework for audio‑visual speech enhancement that uses a large language model to generate natural‑language feedback, which is convert…
Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework
Hsiang-Cheng Yang, You-Jin Li, Rong Chao +3
This paper presents a Gaze-Guided Audio-Visual Speech Enhancement (GG-AVSE) framework to address the cocktail party problem. A major challenge in conventional AVSE is identifying t…
End-to-end audio-visual learning for cochlear implant sound coding simulations in noisy environments
Meng-Ping Lin, Enoch Hsin-Ho Huang, Shao-Yi Chien +1
The cochlear implant (CI) is a successful biomedical device that enables individuals with severe-to-profound hearing loss to perceive sound through electrical stimulation, yet list…
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
Meng-Ping Lin, Jen-Cheng Hou, Chia-Wei Chen +4
Speech enhancement (SE) aims to improve the quality and intelligibility of speech in noisy environments. Recent studies have shown that incorporating visual cues in audio signal pr…