collaborators

7 papers

cs.CV2026

Out of Sight, Still in Mind: Token Compression for Omni-LLMs

Suho Yoo, Youngjoon Jang, Hyebin Cho +1

The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly over audio, video and text, but…

cs.SD2026

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs

Hyebin Cho, Suho Yoo, Jaehyuk Jang +2

While Audio Large Language Models (Audio LLMs) excel at multimodal understanding, they suffer from text dominance, a bias where models blindly favor text over acoustic evidence, ca…

cs.CV2026

On the Nature of Attention Sink that Shapes Decoding Strategy in Omni-LLMs

Suho Yoo, Youngjoon Jang, Joon Son Chung

The goal of this paper is to strengthen the reasoning of Omnimodal Large Language Models (Omni-LLMs) at inference time, without additional training. These models jointly process vi…

cs.CL2026

Speculative End-Turn Detector for Efficient Speech Chatbot Assistant

Hyunjong Ok, Suho Yoo, Jaeho Lee

Spoken dialogue systems powered by large language models have demonstrated remarkable abilities in understanding human speech and generating appropriate spoken responses. However,…

cs.CL2026

AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?

Hyunjong Ok, Suho Yoo, Hyeonjun Kim +1

Even without directly hearing sounds, humans can effortlessly reason about auditory properties, such as pitch, loudness, or sound-source associations, drawing on auditory commonsen…

cs.CL2025

Imagine to Hear: Auditory Knowledge Generation can be an Effective Assistant for Language Models

Suho Yoo, Hyunjong Ok, Jaeho Lee

Language models pretrained on text-only corpora often struggle with tasks that require auditory commonsense knowledge. Previous work addresses this problem by augmenting the langua…