9 papers
Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification
Ming-Hsiang Hu, Kuan-Tang Huang, Chien-Chun Wang +2
User-defined keyword spotting (UD-KWS) enables zero-shot wake-word detection from text, but existing systems learn speaker-invariant representations that cannot reject impostors ut…
Generalized Stock Price Prediction for Multiple Stocks Combined with News Fusion
Pei-Jun Liao, Hung-Shin Lee, Yao-Fei Cheng +3
Predicting stock prices presents challenges in financial forecasting. While traditional approaches such as ARIMA and RNNs are prevalent, recent developments in Large Language Model…
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
Cheng-Yeh Yang, Chien-Chun Wang, Li-Wei Chen +3
Low-resource automatic speech recognition (ASR) continues to pose significant challenges, primarily due to the limited availability of transcribed data for numerous languages. Whil…
SCENE: Semantic-aware Codec Enhancement with Neural Embeddings
Han-Yu Lin, Li-Wei Chen, Hung-Shin Lee
Compression artifacts from standard video codecs often degrade perceptual quality. We propose a lightweight, semantic-aware pre-processing framework that enhances perceptual fideli…
Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages
Yao-Fei Cheng, Li-Wei Chen, Hung-Shin Lee +1
This study investigates the efficacy of data augmentation techniques for low-resource automatic speech recognition (ASR), focusing on two endangered Austronesian languages, Amis an…
DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction
Cheng-Yeh Yang, Kuan-Tang Huang, Chien-Chun Wang +3
A pooling mechanism is essential for mean opinion score (MOS) prediction, facilitating the transformation of variable-length audio features into a concise fixed-size representation…