most citedEfficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer

2 citations · 2 across the 4 of their papers we have counts for

collaborators

6 papers

cs.SD2026

Exploiting Speech LLM Representations for Multilingual and Cross-Lingual Parkinson's Disease Detection

Sarthak Giri, Zi Haur Pang, Tatsuya Kawahara

Speech Large Language Models (Speech LLMs) have shown strong performance across diverse tasks, yet their utility for pathological speech analysis remains underexplored. In this wor…

cs.MM2026

Learning to Prefer Reliably: Error-Augmented Emotion Preference Optimization with Calibrated Fusion

Zilong Huang, Junyi Peng, Junjie Li +5

Emotion preference learning uses pairwise comparisons between candidate descriptions to align multimodal large language models (MLLMs) with human judgments of open-ended emotion de…

cs.SD2026

Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features

Ryo Magoshi, Jaeyoung Lee, Shinsuke Sakai +1

Recent Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription rely on Grapheme-to-Phoneme (G2P) labels, but the phoneme labels are not necessarily phonetically faithful.…

cs.CL2026

LoASR-Bench: Evaluating Large Speech Language Models on Low-Resource Automatic Speech Recognition Across Language Families

Jianan Chen, Xiaoxue Gao, Tatsuya Kawahara +1

Large language models (LLMs) have driven substantial advances in speech language models (SpeechLMs), yielding strong performance in automatic speech recognition (ASR) under high-re…

cs.SD2026

ERM-MinMaxGAP: Benchmarking and Mitigating Gender Bias in Multilingual Multimodal Speech-LLM Emotion Recognition

Zi Haur Pang, Xiaoxue Gao, Tatsuya Kawahara +1

Speech emotion recognition (SER) systems can exhibit gender-related performance disparities, but how such bias manifests in multilingual speech LLMs across languages and modalities…

cs.SD2024★ 2 cited

Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer

Tomoki Honda, Shinsuke Sakai, Tatsuya Kawahara

Recently, Conformer has achieved state-of-the-art performance in many speech recognition tasks. However, the Transformer-based models show significant deterioration for long-form s…