From the 1 of 12 linked papers with an AI index.
12 papers
The SonicAGI System for the REAL-TSE Challenge
Kai Li, Wendi Sang, Jintao Cheng +1
The paper presents SonicAGI, a system for real-world target speaker extraction that combines simulated and real meeting data, using a low‑latency SwiftNet-Lookahead model for onlin…
A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
Kai Li, Jintao Cheng, Chang Zeng +5
Query-based universal sound separation is fundamental to intelligent auditory systems, aiming to isolate specific sources from mixtures. Despite recent advances, existing methods c…
A Brain-Inspired Deep Separation Network for Single Channel Raman Spectra Unmixing
Gaoruishu Long, Jinchao Liu, Bo Liu +2
Raman spectra obtained in real world applications are often a noisy combination of several spectra of various substances in a tested sample. Unmixing such spectra into individual c…
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
Kai Li, Kejun Gao, Xiaolin Hu
Audio-visual speech separation (AVSS) methods leverage visual cues to extract target speech and have demonstrated strong separation quality in noisy acoustic environments. However,…
TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation
Mohan Xu, Kai Li, Guo Chen +1
In recent years, much speech separation research has focused primarily on improving model performance. However, for low-latency speech processing systems, high efficiency is equall…
A-LLM: An End-to-end Conversational Audio Avatar Large Language Model
Xiaolin Hu, Hang Yuan, Xinzhu Sang +4
Developing expressive and responsive conversational digital humans is a cornerstone of next-generation human-computer interaction. While large language models (LLMs) have significa…