2 citations · 3 across the 4 of their papers we have counts for
Showing 2025Show all
2 papers · 1 filter
cs.SD2025
FCPE: A Fast Context-based Pitch Estimation Model
Yuxin Luo, Ruoyi Zhang, Lu-Chuan Liu +2
Pitch estimation (PE) in monophonic audio is crucial for MIDI transcription and singing voice conversion (SVC), but existing methods suffer significant performance degradation unde…
cs.SD2025★ 2 cited
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling
Yifan Cheng, Ruoyi Zhang, Jiatong Shi
Acquiring large-scale emotional speech data with strong consistency remains a challenge for speech synthesis. This paper presents MIKU-PAL, a fully automated multimodal pipeline fo…