2 papers
cs.CV2026
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
Hyeonwoo Cho, Donghyeon Baek, Yewon Kim +1
Recent advancements in Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks, yet the quadratic computational complexity arising from t…
cs.CV2024
SoundBrush: Sound as a Brush for Visual Scene Editing
Kim Sung-Bin, Kim Jun-Seong, Junseok Ko +2
We propose SoundBrush, a model that uses sound as a brush to edit and manipulate visual scenes. We extend the generative capabilities of the Latent Diffusion Model (LDM) to incorpo…