collaborators

13 papers

cs.RO2026

Casting Everything to Online API Services? A Survey of Integrating Localized Speech Recognition Models in Robotic Systems

Sheng Li, Jing Li, Felix Schijve +2

The paper surveys how automatic speech recognition models, from classic methods to modern deep‑learning systems like Whisper, are integrated into robotic platforms using online API…

eess.AS2026

Synchronized Three-Dimensional Vocal-Tract Motion for Speech Synchronization via Joint-Embedding Predictive Architecture Alignment

Sheng Li, Takahiro Shinozaki

The paper presents a method that synchronizes a three‑dimensional articulatory vocal‑tract model with an acoustic carrier waveform using a joint‑embedding predictive architecture a…

cs.CL2026

Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition

Zhengdong Yang, Qianying Liu, Sheng Li +2

We present a novel approach centered on the decoding stage of Automatic Speech Recognition (ASR) that enhances multilingual performance, especially for low-resource languages. It u…

cs.CL2025

SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models

Zhen Wan, Chao-Han Huck Yang, Yahan Yu +8

We introduce Speech-based Intelligence Quotient (SIQ) as a new form of human cognition-inspired evaluation pipeline for voice understanding large language models, LLM Voice, design…

cs.SD2025

Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement

Jianing Yang, Sheng Li, Takahiro Shinozaki +2

Current emotional Text-To-Speech (TTS) and style transfer methods rely on reference encoders to control global style or emotion vectors, but do not capture nuanced acoustic details…

eess.AS2025

RAG-Boost: Retrieval-Augmented Generation Enhanced LLM-based Speech Recognition

Pengcheng Wang, Sheng Li, Takahiro Shinozaki

In this paper, we propose RAG-Boost (ST-ShinozakiLab Task I system), which enhances the baseline LLM-based ASR system of the MLC-SLM Challenge (task I) with a retrieval-augmented g…