activity
20242026
collaborators

9 papers

cs.SD2026

Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

Oh Hyun-Bin, Kazuki Shimada, Yuhta Takida +6

Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, s…

cs.CL2026

A Language-Guided Bayesian Optimization for Efficient LoRA Hyperparameter Search

Baek Seong-Eun, Lee Jung-Mok, Kim Sung-Bin +1

Fine-tuning Large Language Models (LLMs) with Low-Rank Adaptation (LoRA) offers a resource-efficient way to personalize or specialize. However, LoRA is highly sensitive to hyperpar…

cs.CL2026

SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter

Lee Jung-Mok, Kim Sung-Bin, Joohyun Chang +2

Laughter is a complex social signal that conveys communicative intent beyond amusement. While prior work has focused on isolated laughter analysis tasks, a comprehensive understand…

cs.CV2025

FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling

Kim Sung-Bin, Joohyun Chang, David Harwath +1

Talking face editing and face generation have often been studied as distinct problems. In this work, we propose viewing both not as separate tasks but as subtasks of a unifying for…

cs.GR2025

Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation Metrics

Lee Chae-Yeon, Oh Hyun-Bin, Han EunGi +3

Recent advancements in speech-driven 3D talking head generation have made significant progress in lip synchronization. However, existing models still struggle to capture the percep…

cs.CV2024

SoundBrush: Sound as a Brush for Visual Scene Editing

Kim Sung-Bin, Kim Jun-Seong, Junseok Ko +2

We propose SoundBrush, a model that uses sound as a brush to edit and manipulate visual scenes. We extend the generative capabilities of the Latent Diffusion Model (LDM) to incorpo…