11 papers
Out of Sight, Still in Mind: Token Compression for Omni-LLMs
Suho Yoo, Youngjoon Jang, Hyebin Cho +1
The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly over audio, video and text, but…
See & Sniff: Learning Visuo-Olfactory Representations
Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu +2
While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olfactory data. We introduce Sm…
ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion
Jeongsoo Choi, Ji-Hoon Kim, Shujie Hu +1
Neural speech codecs efficiently compress speech and have become a foundation for speech generation, but they are typically learned as holistic representations that intertwine ling…
Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs
Hyebin Cho, Suho Yoo, Jaehyuk Jang +2
While Audio Large Language Models (Audio LLMs) excel at multimodal understanding, they suffer from text dominance, a bias where models blindly favor text over acoustic evidence, ca…
Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models
Hyebin Cho, Jaehyuk Jang, Changick Kim +1
Audio-Language Models (ALMs) have shown remarkable success in zero-shot audio classification by aligning audio waveforms with text. Recent efforts to improve downstream performance…
Inference-Time Scaling for Joint Audio-Video Generation
Jaemin Jung, Kyeongha Rho, Inkyu Shin +1
Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchronized. While existing joint au…