5 papers
Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization
Tony Alex, Wish Suharitdamrong, Sara Atito +5
Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content…
The Hidden Evolution of Disguised Visual Context inside the VLM
Wish Suharitdamrong, Tony Alex, Xiatian Zhu +2
Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and interact with the language space depends enti…
CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks
Wish Suharitdamrong, Tony Alex, Muhammad Awais +1
Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures composed of unimodal encoders, such as DINO…
SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes
Tony Alex, Sara Ahmed, Armin Mustafa +2
Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed…
PAL: Probing Audio Encoders via LLMs -- Audio Information Transfer into LLMs
Tony Alex, Wish Suharitdamrong, Sara Atito +4
Integration of audio perception into large language models (LLMs) is an emerging research area for enabling machine listening applications, yet efficient transfer of rich audio sem…