activity
20242026
collaborators

8 papers

cs.SD2026

AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models

Junqi Zhao, Chenxing Li, Jinzheng Zhao +4

We propose a general feedback-driven retrieval-augmented generation (RAG) approach that leverages Large Audio Language Models (LALMs) to address the missing or imperfect synthesis…

cs.CL2025

Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory Access

Xiang Hu, Jiaqi Leng, Jun Zhao +2

A key advantage of Recurrent Neural Networks (RNNs) over Transformers is their linear computational and space complexity enables faster training and inference for long sequences. H…

eess.AS2025

Region-Specific Audio Tagging for Spatial Sound

Jinzheng Zhao, Yong Xu, Haohe Liu +6

Audio tagging aims to label sound events appearing in an audio recording. In this paper, we propose region-specific audio tagging, a new task which labels sound events in a given r…

eess.AS2025

WTFormer: A Wavelet Conformer Network for MIMO Speech Enhancement with Spatial Cues Peservation

Lu Han, Junqi Zhao, Renhua Peng

Current multi-channel speech enhancement systems mainly adopt single-output architecture, which face significant challenges in preserving spatio-temporal signal integrity during mu…

cs.CL2025

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English

Haoyang Zhang, Hexin Liu, Xiangyu Zhang +7

The speech tokenizer plays a crucial role in recent speech tasks, generally serving as a bridge between speech signals and language models. While low-frame-rate codecs are widely e…

cs.SD2025

AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion

Junqi Zhao, Jinzheng Zhao, Haohe Liu +5

Diffusion models have significantly improved the quality and diversity of audio generation but are hindered by slow inference speed. Rectified flow enhances inference speed by lear…