98 citations · 133 across the 2 of their papers we have counts for
4 papers
Controllable Neural Prosody Synthesis
Max Morrison, Zeyu Jin, Justin Salamon +2
Speech synthesis has recently seen significant improvements in fidelity, driven by the advent of neural vocoders and neural prosody generators. However, these systems lack intuitiv…
F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder
Kaizhi Qian, Zeyu Jin, Mark Hasegawa-Johnson +1
Non-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Many style-transfer-inspired methods such as generative adversarial networ…
A Differentiable Perceptual Audio Metric Learned from Just Noticeable Differences
Pranay Manocha, Adam Finkelstein, Richard Zhang +3
Many audio processing tasks require perceptual assessment. The ``gold standard`` of obtaining human judgments is time-consuming, expensive, and cannot be used as an optimization cr…
B-Script: Transcript-based B-roll Video Editing with Recommendations
Bernd Huber, Hijung Valentina Shin, Bryan Russell +2
In video production, inserting B-roll is a widely used technique to enrich the story and make a video more engaging. However, determining the right content and positions of B-roll…