4 papers
Inference-Time Scaling for Joint Audio-Video Generation
Jaemin Jung, Kyeongha Rho, Inkyu Shin +1
Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchronized. While existing joint au…
Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs
Chaeyoung Jung, Kyeongha Rho, Joon Son Chung
Omnimodal Large Language Models (Omni-LLMs) incur substantial computational overhead due to the large number of multimodal input tokens they process, making token reduction essenti…
MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
Kyeongha Rho, Hyeongkeun Lee, Jae Won Cho +1
In this paper, we propose Mixture of Layer-Wise Tokens (MoLT), a parameter- and memory-efficient adaptation framework for audio-visual learning. The key idea of MoLT is to replace…
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
Kyeongha Rho, Hyeongkeun Lee, Valentio Iverson +1
Automated audio captioning is a task that generates textual descriptions for audio content, and recent studies have explored using visual information to enhance captioning quality.…