4 papers
Preference Optimization with LALM Feedback for Continuous Autoregressive Non-Verbal Vocalization Generation
Jingbin Hu, Qirui Zhan, Yuang Cao +7
We propose a preference optimization framework with Large Audio-Language Model (LALM) feedback for controllable non-verbal vocalization (NVV) generation in continuous autoregressiv…
SpeechAnnotator: A Context-Aware Multi-Agent Framework and Benchmark for Multidimensional Speech Annotation
Qirui Zhan, Shuiyuan Wang, Jingbin Hu +10
Recent controllable speech generation requires training data with fine-grained annotations of speaker traits, prosody, emotion, paralinguistic cues, acoustic scenes, and context. E…
Source-Adaptive Data Curation for Bilingual NVV-Aware ASR
Yuang Cao, Qirui Zhan, Jingbin Hu +7
Nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, convey affective and interactional information that conventional automatic speech recognition (ASR) sy…
SphereVAE: Hyperspherical Latent Autoencoders for Robust Autoregressive Speech Representation Modeling
Haoyu Zhang, Jingbin Hu, Hanke Xie +8
With the rapid development of speech generation technology, discrete codec representations have been widely used because they provide a stable prediction paradigm. In expressive sp…