4 papers
Hierarchical Activity Recognition and Captioning from Long-Form Audio
Peng Zhang, Qingyu Luo, Philip J. B. Jackson +1
Complex activities in real-world audio unfold over extended durations and exhibit hierarchical structure, yet most prior work focuses on short clips and isolated events. To bridge…
ISSE: An Instruction-Guided Speech Style Editing Dataset And Benchmark
Yun Chen, Qi Chen, Zheqi Dai +3
Speech style editing refers to modifying the stylistic properties of speech while preserving its linguistic content and speaker identity. However, most existing approaches depend o…
SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes
Tony Alex, Sara Ahmed, Armin Mustafa +2
Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed…
PAL: Probing Audio Encoders via LLMs -- Audio Information Transfer into LLMs
Tony Alex, Wish Suharitdamrong, Sara Atito +4
Integration of audio perception into large language models (LLMs) is an emerging research area for enabling machine listening applications, yet efficient transfer of rich audio sem…