5 papers
Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources
Oh Hyun-Bin, Kazuki Shimada, Yuhta Takida +6
Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, s…
A Language-Guided Bayesian Optimization for Efficient LoRA Hyperparameter Search
Baek Seong-Eun, Lee Jung-Mok, Kim Sung-Bin +1
Fine-tuning Large Language Models (LLMs) with Low-Rank Adaptation (LoRA) offers a resource-efficient way to personalize or specialize. However, LoRA is highly sensitive to hyperpar…
FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling
Kim Sung-Bin, Joohyun Chang, David Harwath +1
Talking face editing and face generation have often been studied as distinct problems. In this work, we propose viewing both not as separate tasks but as subtasks of a unifying for…
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
Jeongsoo Choi, Ji-Hoon Kim, Kim Sung-Bin +2
In this paper, we address the task of multimodal-to-speech generation, which aims to synthesize high-quality speech from multiple input modalities: text, video, and reference audio…
SoundBrush: Sound as a Brush for Visual Scene Editing
Kim Sung-Bin, Kim Jun-Seong, Junseok Ko +2
We propose SoundBrush, a model that uses sound as a brush to edit and manipulate visual scenes. We extend the generative capabilities of the Latent Diffusion Model (LDM) to incorpo…