activity
20232025
collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2025

SpeechMLC: Speech Multi-label Classification

Miseul Kim, Seyun Um, Hyeonjin Cha +1

In this paper, we propose a multi-label classification framework to detect multiple speaking styles in a speech sample. Unlike previous studies that have primarily focused on ident…

eess.AS2025

Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation

Miseul Kim, Soo Jin Park, Kyungguen Byun +4

Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speak…

eess.AS2024

Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation

Miseul Kim, Soo-Whan Chung, Youna Ji +2

This paper introduces a novel task in generative speech processing, Acoustic Scene Transfer (AST), which aims to transfer acoustic scenes of speech signals to diverse environments.…

eess.AS2023

Self-supervised Complex Network for Machine Sound Anomaly Detection

Miseul Kim, Minh Tri Ho, Hong-Goo Kang

In this paper, we propose an anomaly detection algorithm for machine sounds with a deep complex network trained by self-supervision. Using the fact that phase continuity informatio…

eess.AS2023

Style Modeling for Multi-Speaker Articulation-to-Speech

Miseul Kim, Zhenyu Piao, Jihyun Lee +1

In this paper, we propose a neural articulation-to-speech (ATS) framework that synthesizes high-quality speech from articulatory signal in a multi-speaker situation. Most conventio…

eess.AS2023

BrainTalker: Low-Resource Brain-to-Speech Synthesis with Transfer Learning using Wav2Vec 2.0

Miseul Kim, Zhenyu Piao, Jihyun Lee +1

Decoding spoken speech from neural activity in the brain is a fast-emerging research topic, as it could enable communication for people who have difficulties with producing audible…