collaborators

5 papers

cs.SD2026

Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization

Tony Alex, Wish Suharitdamrong, Sara Atito +5

Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content…

cs.SD2026

Hierarchical Activity Recognition and Captioning from Long-Form Audio

Peng Zhang, Qingyu Luo, Philip J. B. Jackson +1

Complex activities in real-world audio unfold over extended durations and exhibit hierarchical structure, yet most prior work focuses on short clips and isolated events. To bridge…

cs.SD2026

PAL: Probing Audio Encoders via LLMs -- Audio Information Transfer into LLMs

Tony Alex, Wish Suharitdamrong, Sara Atito +4

Integration of audio perception into large language models (LLMs) is an emerging research area for enabling machine listening applications, yet efficient transfer of rich audio sem…

eess.AS2025

ISSE: An Instruction-Guided Speech Style Editing Dataset And Benchmark

Yun Chen, Qi Chen, Zheqi Dai +3

Speech style editing refers to modifying the stylistic properties of speech while preserving its linguistic content and speaker identity. However, most existing approaches depend o…

cs.SD2025

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

Tony Alex, Sara Ahmed, Armin Mustafa +2

Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed…