6 papers
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis
Zhisheng Zheng, Xiaohang Sun, Zhu Liu +5
Recent Text-To-Speech (TTS) systems have achieved strong naturalness and zero-shot voice cloning performance, but fine-grained control of expressive speech at the word or phoneme l…
AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation
Dongmei Wang, Xiaohang Sun, Yang Liu +10
We propose AugCodec, a low-bitrate disentangled neural speech codec that leverages data augmentation to decompose speech into three distinct components: semantic, speaker, and pros…
From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer Decomposition
Jingxi Chen, Yixiao Zhang, Xiaoye Qian +4
Images can be viewed as layered compositions, foreground objects over background, with potential occlusions. This layered representation enables independent editing of elements, of…
Subtle Motion Blur Detection and Segmentation from Static Image Artworks
Ganesh Samarth, Sibendu Paul, Solale Tabarestani +1
Streaming services serve hundreds of millions of viewers worldwide, where visual assets such as thumbnails, box art, and cover images are critical for engagement. Subtle motion blu…
Automatic Funny Scene Extraction from Long-form Cinematic Videos
Sibendu Paul, Haotian Jiang, Caren Chen
Automatically extracting engaging and high-quality humorous scenes from cinematic titles is pivotal for creating captivating video previews and snackable content, boosting user eng…
From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
Shih-Yao Lin, Sibendu Paul, Caren Chen
Selecting informative keyframes is critical for efficient video understanding, yet existing approaches often rely on heuristics, ignore semantics, or produce redundant frames. We p…