collaborators

5 papers

cs.SD2026

Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization

Tony Alex, Wish Suharitdamrong, Sara Atito +5

Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content…

cs.AI2026

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs

Wish Suharitdamrong, Muhammad Awais, Xiatian Zhu +1

Multimodal Large Language Models (MLLMs) can listen and see, but how do audio and visual signals actually travel through the network to shape an answer? Despite their growing role…

cs.CV2026

The Hidden Evolution of Disguised Visual Context inside the VLM

Wish Suharitdamrong, Tony Alex, Xiatian Zhu +2

Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and interact with the language space depends enti…

cs.CV2026

CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

Wish Suharitdamrong, Tony Alex, Muhammad Awais +1

Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures composed of unimodal encoders, such as DINO…

cs.SD2025

PAL: Probing Audio Encoders via LLMs -- Audio Information Transfer into LLMs

Tony Alex, Wish Suharitdamrong, Sara Atito +4

Integration of audio perception into large language models (LLMs) is an emerging research area for enabling machine listening applications, yet efficient transfer of rich audio sem…