13 papers
Casting Everything to Online API Services? A Survey of Integrating Localized Speech Recognition Models in Robotic Systems
Sheng Li, Jing Li, Felix Schijve +2
The paper surveys how automatic speech recognition models, from classic methods to modern deep‑learning systems like Whisper, are integrated into robotic platforms using online API…
Synchronized Three-Dimensional Vocal-Tract Motion for Speech Synchronization via Joint-Embedding Predictive Architecture Alignment
Sheng Li, Takahiro Shinozaki
The paper presents a method that synchronizes a three‑dimensional articulatory vocal‑tract model with an acoustic carrier waveform using a joint‑embedding predictive architecture a…
Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition
Zhengdong Yang, Qianying Liu, Sheng Li +2
We present a novel approach centered on the decoding stage of Automatic Speech Recognition (ASR) that enhances multilingual performance, especially for low-resource languages. It u…
SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models
Zhen Wan, Chao-Han Huck Yang, Yahan Yu +8
We introduce Speech-based Intelligence Quotient (SIQ) as a new form of human cognition-inspired evaluation pipeline for voice understanding large language models, LLM Voice, design…
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
Jianing Yang, Sheng Li, Takahiro Shinozaki +2
Current emotional Text-To-Speech (TTS) and style transfer methods rely on reference encoders to control global style or emotion vectors, but do not capture nuanced acoustic details…
RAG-Boost: Retrieval-Augmented Generation Enhanced LLM-based Speech Recognition
Pengcheng Wang, Sheng Li, Takahiro Shinozaki
In this paper, we propose RAG-Boost (ST-ShinozakiLab Task I system), which enhances the baseline LLM-based ASR system of the MLC-SLM Challenge (task I) with a retrieval-augmented g…