6 papers
Efficient Punctuation Restoration via Weighted Lookahead Scoring Method for Streaming ASR Systems
Sungmook Woo, Hyungu Kang, Chanwoo Kim
Punctuation restoration improves ASR (Automatic Speech Recognition) readability. However streaming ASR requires online decisions with limited future context. In streaming ASR, the…
HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
Gio Paik, Yongbeom Kim, Soungmin Lee +2
Despite advances in multilingual automatic speech recognition (ASR), code-switching (CS), the mixing of languages within an utterance common in daily speech, remains a severely und…
Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence
Yerin Ryu, Inseop Shin, Chanwoo Kim
Controllable Singing Voice Synthesis (SVS) aims to generate expressive singing voices reflecting user intent. While recent SVS systems achieve high audio quality, most rely on prob…
OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics
Wei Chu, Yuanzhe Dong, Ke Tan +7
OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podca…
HyperFlow: Gradient-Free Emulation of Few-Shot Fine-Tuning
Donggyun Kim, Chanwoo Kim, Seunghoon Hong
While test-time fine-tuning is beneficial in few-shot learning, the need for multiple backpropagation steps can be prohibitively expensive in real-time or low-resource scenarios. T…
AV-Surf: Surface-Enhanced Geometry-Aware Novel-View Acoustic Synthesis
Hadam Baek, Hannie Shin, Jiyoung Seo +4
Accurately modeling sound propagation with complex real-world environments is essential for Novel View Acoustic Synthesis (NVAS). While previous studies have leveraged visual perce…