2 citations · 2 across the 4 of their papers we have counts for
6 papers
Exploiting Speech LLM Representations for Multilingual and Cross-Lingual Parkinson's Disease Detection
Sarthak Giri, Zi Haur Pang, Tatsuya Kawahara
Speech Large Language Models (Speech LLMs) have shown strong performance across diverse tasks, yet their utility for pathological speech analysis remains underexplored. In this wor…
Learning to Prefer Reliably: Error-Augmented Emotion Preference Optimization with Calibrated Fusion
Zilong Huang, Junyi Peng, Junjie Li +5
Emotion preference learning uses pairwise comparisons between candidate descriptions to align multimodal large language models (MLLMs) with human judgments of open-ended emotion de…
Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features
Ryo Magoshi, Jaeyoung Lee, Shinsuke Sakai +1
Recent Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription rely on Grapheme-to-Phoneme (G2P) labels, but the phoneme labels are not necessarily phonetically faithful.…
LoASR-Bench: Evaluating Large Speech Language Models on Low-Resource Automatic Speech Recognition Across Language Families
Jianan Chen, Xiaoxue Gao, Tatsuya Kawahara +1
Large language models (LLMs) have driven substantial advances in speech language models (SpeechLMs), yielding strong performance in automatic speech recognition (ASR) under high-re…
ERM-MinMaxGAP: Benchmarking and Mitigating Gender Bias in Multilingual Multimodal Speech-LLM Emotion Recognition
Zi Haur Pang, Xiaoxue Gao, Tatsuya Kawahara +1
Speech emotion recognition (SER) systems can exhibit gender-related performance disparities, but how such bias manifests in multilingual speech LLMs across languages and modalities…
Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer
Tomoki Honda, Shinsuke Sakai, Tatsuya Kawahara
Recently, Conformer has achieved state-of-the-art performance in many speech recognition tasks. However, the Transformer-based models show significant deterioration for long-form s…