3 papers
cs.AI2026
Relative Time Intervals Representation for Word-level Timestamping with Masked Training
Quanwei Tang, Zhiyu Tang, Xu Li +3
Although Speech Large Language Models (SpeechLLMs) excel at speech understanding and generation, their capacity for fine-grained, temporally aligned outputs remains underexplored.…
cs.SD2024
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
Xiaoyu Liu, Xu Li, Joan Serrà +1
Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative mode…
cs.SD2024
MaskSR: Masked Language Model for Full-band Speech Restoration
Xu Li, Qirui Wang, Xiaoyu Liu
Speech restoration aims at restoring high quality speech in the presence of a diverse set of distortions. Although several deep learning paradigms have been studied for this task,…