From the 1 of 19 linked papers with an AI index.
19 papers
Efficient Text-to-Audio Generation via Pruning
Arshdeep Singh, Yi Yuan, Yun Chen +2
The paper applies filter‑based pruning to the U‑Net backbone of the AudioLDM text‑to‑audio diffusion model, reducing most of its parameters and compute while preserving generation…
Explainable AI in Speaker Recognition -- Making Latent Representations Understandable
Yanze Xu, Wenwu Wang, Mark D. Plumbley
Neural networks can be trained to learn task-relevant representations from data. Understanding how these networks make decisions falls within the Explainable AI (XAI) domain. This…
Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation
Yanze Xu, Mark D. Plumbley, Wenwu Wang
Explaining and understanding the decision-making process of artificial intelligence (AI) systems, particularly those implemented by neural networks, falls within the field of expla…
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
Yi Yuan, Xubo Liu, Haohe Liu +5
With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite produc…
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
Junqi Zhao, Chenxing Li, Jinzheng Zhao +4
We propose a general feedback-driven retrieval-augmented generation (RAG) approach that leverages Large Audio Language Models (LALMs) to address the missing or imperfect synthesis…
EnvSDD: Benchmarking Environmental Sound Deepfake Detection
Han Yin, Yang Xiao, Rohan Kumar Das +4
Audio generation systems now create very realistic soundscapes that can enhance media production, but also pose potential risks. Several studies have examined deepfakes in speech o…