activity
20242026
collaborators

6 papers

cs.SD2026

MixProLAP: Mixture-Induced Uncertainty Modeling for Probabilistic Language-Audio Pretraining

Yu Nakagome, Jaesong Lee, Soo-Whan Chung

Acoustic environments often contain multiple overlapping sound events, and the same acoustic scene can be described using diverse textual expressions, making audio-text alignment i…

eess.AS2025

Seeing What You Say: Expressive Image Generation from Speech

Jiyoung Lee, Song Park, Sanghyuk Chun +1

This paper proposes VoxStudio, the first unified and end-to-end speech-to-image model that generates expressive images directly from spoken descriptions by jointly aligning linguis…

eess.AS2025

MF-PAM: Accurate Pitch Estimation through Periodicity Analysis and Multi-level Feature Fusion

Woo-Jin Chung, Doyeon Kim, Soo-Whan Chung +1

We introduce Multi-level feature Fusion-based Periodicity Analysis Model (MF-PAM), a novel deep learning-based pitch estimation model that accurately estimates pitch trajectory in…

eess.AS2025

Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation

Soo-Whan Chung, Min-Seok Choi

This paper introduces a novel approach to speech restoration by integrating a context-related conditioning strategy. Specifically, we employ the diffusion-based generative restorat…

cs.CV2025

S3D: Sketch-Driven 3D Model Generation

Hail Song, Wonsik Shin, Naeun Lee +3

Generating high-quality 3D models from 2D sketches is a challenging task due to the inherent ambiguity and sparsity of sketch data. In this paper, we present S3D, a novel framework…

eess.AS2024

Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation

Miseul Kim, Soo-Whan Chung, Youna Ji +2

This paper introduces a novel task in generative speech processing, Acoustic Scene Transfer (AST), which aims to transfer acoustic scenes of speech signals to diverse environments.…