3 papers
cs.CV2026
HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding
Wenhui Liao, Hongliang Li, Pengyu Xie +15
Document parsing is a fundamental task in multimodal understanding, supporting a wide range of downstream applications such as information extraction and intelligent document analy…
eess.AS2026
VoiceSculptor: Your Voice, Designed By You
Jingbin Hu, Huakang Chen, Linhan Ma +19
Despite rapid progress in text-to-speech (TTS), open-source systems still lack truly instruction-following, fine-grained control over core speech attributes (e.g., pitch, speaking…
cs.SD2025
NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations
Huan Liao, Qinke Ni, Yuancheng Wang +5
Paralinguistic vocalizations-including non-verbal sounds like laughter and breathing, as well as lexicalized interjections such as "uhm" and "oh"-are integral to natural spoken com…