7 papers · 1 filter
SpeechMLC: Speech Multi-label Classification
Miseul Kim, Seyun Um, Hyeonjin Cha +1
In this paper, we propose a multi-label classification framework to detect multiple speaking styles in a speech sample. Unlike previous studies that have primarily focused on ident…
Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
Miseul Kim, Soo Jin Park, Kyungguen Byun +4
Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speak…
Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation
Miseul Kim, Soo-Whan Chung, Youna Ji +2
This paper introduces a novel task in generative speech processing, Acoustic Scene Transfer (AST), which aims to transfer acoustic scenes of speech signals to diverse environments.…
Self-supervised Complex Network for Machine Sound Anomaly Detection
Miseul Kim, Minh Tri Ho, Hong-Goo Kang
In this paper, we propose an anomaly detection algorithm for machine sounds with a deep complex network trained by self-supervision. Using the fact that phase continuity informatio…
Style Modeling for Multi-Speaker Articulation-to-Speech
Miseul Kim, Zhenyu Piao, Jihyun Lee +1
In this paper, we propose a neural articulation-to-speech (ATS) framework that synthesizes high-quality speech from articulatory signal in a multi-speaker situation. Most conventio…
BrainTalker: Low-Resource Brain-to-Speech Synthesis with Transfer Learning using Wav2Vec 2.0
Miseul Kim, Zhenyu Piao, Jihyun Lee +1
Decoding spoken speech from neural activity in the brain is a fast-emerging research topic, as it could enable communication for people who have difficulties with producing audible…