collaborators

11 papers

cs.CV2026

Out of Sight, Still in Mind: Token Compression for Omni-LLMs

Suho Yoo, Youngjoon Jang, Hyebin Cho +1

The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly over audio, video and text, but…

cs.CV2026

See & Sniff: Learning Visuo-Olfactory Representations

Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu +2

While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olfactory data. We introduce Sm…

eess.AS2026

ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion

Jeongsoo Choi, Ji-Hoon Kim, Shujie Hu +1

Neural speech codecs efficiently compress speech and have become a foundation for speech generation, but they are typically learned as holistic representations that intertwine ling…

cs.SD2026

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs

Hyebin Cho, Suho Yoo, Jaehyuk Jang +2

While Audio Large Language Models (Audio LLMs) excel at multimodal understanding, they suffer from text dominance, a bias where models blindly favor text over acoustic evidence, ca…

cs.SD2026

Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models

Hyebin Cho, Jaehyuk Jang, Changick Kim +1

Audio-Language Models (ALMs) have shown remarkable success in zero-shot audio classification by aligning audio waveforms with text. Recent efforts to improve downstream performance…

cs.MM2026

Inference-Time Scaling for Joint Audio-Video Generation

Jaemin Jung, Kyeongha Rho, Inkyu Shin +1

Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchronized. While existing joint au…