works on

From the 1 of 19 linked papers with an AI index.

activity
20242026
collaborators

19 papers

eess.AS2026

Efficient Text-to-Audio Generation via Pruning

Arshdeep Singh, Yi Yuan, Yun Chen +2

The paper applies filter‑based pruning to the U‑Net backbone of the AudioLDM text‑to‑audio diffusion model, reducing most of its parameters and compute while preserving generation…

eess.AS2026

Explainable AI in Speaker Recognition -- Making Latent Representations Understandable

Yanze Xu, Wenwu Wang, Mark D. Plumbley

Neural networks can be trained to learn task-relevant representations from data. Understanding how these networks make decisions falls within the Explainable AI (XAI) domain. This…

eess.AS2026

Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation

Yanze Xu, Mark D. Plumbley, Wenwu Wang

Explaining and understanding the decision-making process of artificial intelligence (AI) systems, particularly those implemented by neural networks, falls within the field of expla…

cs.SD2026

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models

Yi Yuan, Xubo Liu, Haohe Liu +5

With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite produc…

cs.SD2026

AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models

Junqi Zhao, Chenxing Li, Jinzheng Zhao +4

We propose a general feedback-driven retrieval-augmented generation (RAG) approach that leverages Large Audio Language Models (LALMs) to address the missing or imperfect synthesis…

cs.SD2025

EnvSDD: Benchmarking Environmental Sound Deepfake Detection

Han Yin, Yang Xiao, Rohan Kumar Das +4

Audio generation systems now create very realistic soundscapes that can enhance media production, but also pose potential risks. Several studies have examined deepfakes in speech o…