4 papers
Efficient Text-to-Audio Generation via Pruning
Arshdeep Singh, Yi Yuan, Yun Chen +2
The paper applies filter‑based pruning to the U‑Net backbone of the AudioLDM text‑to‑audio diffusion model, reducing most of its parameters and compute while preserving generation…
Compressing Quaternion Convolutional Neural Networks for Audio Classification
Arshdeep Singh, Vinayak Abrol, Mark D. Plumbley
Conventional Convolutional Neural Networks (CNNs) in the real domain have been widely used for audio classification. However, their convolution operations process multi-channel inp…
ISSE: An Instruction-Guided Speech Style Editing Dataset And Benchmark
Yun Chen, Qi Chen, Zheqi Dai +3
Speech style editing refers to modifying the stylistic properties of speech while preserving its linguistic content and speaker identity. However, most existing approaches depend o…
Integrating IP Broadcasting with Audio Tags: Workflow and Challenges
Rhys Burchett-Vass, Arshdeep Singh, Gabriel Bibbó +1
The broadcasting industry has adopted IP technologies, revolutionising both live and pre-recorded content production, from news gathering to live music events. IP broadcasting allo…