6 papers
A-LLM: An End-to-end Conversational Audio Avatar Large Language Model
Xiaolin Hu, Hang Yuan, Xinzhu Sang +4
Developing expressive and responsive conversational digital humans is a cornerstone of next-generation human-computer interaction. While large language models (LLMs) have significa…
FGNet: Leveraging Feature-Guided Attention to Refine SAM2 for 3D EM Neuron Segmentation
Zhenghua Li, Hang Chen, Zihao Sun +2
Accurate segmentation of neural structures in Electron Microscopy (EM) images is paramount for neuroscience. However, this task is challenged by intricate morphologies, low signal-…
Advances in Speech Separation: Techniques, Challenges, and Future Trends
Kai Li, Guo Chen, Wendi Sang +8
The field of speech separation, addressing the "cocktail party problem", has seen revolutionary advances with DNNs. Speech separation enhances clarity in complex acoustic environme…
Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder
Runxuan Yang, Kai Li, Guo Chen +1
This paper addresses the challenge of enhancing the realism of vocoder-generated singing voice audio by mitigating the distinguishable disparities between synthetic and real-life r…
A Fast and Lightweight Model for Causal Audio-Visual Speech Separation
Wendi Sang, Kai Li, Runxuan Yang +2
Audio-visual speech separation (AVSS) aims to extract a target speech signal from a mixed signal by leveraging both auditory and visual (lip movement) cues. However, most existing…
Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech Separation
Guo Chen, Kai Li, Runxuan Yang +1
Existing causal speech separation models often underperform compared to non-causal models due to difficulties in retaining historical information. To address this, we propose the T…