5 papers
X-VC: Zero-shot Streaming Voice Conversion in Codec Space
Qixi Zheng, Yuxiang Zhao, Tianrui Wang +7
Zero-shot voice conversion (VC) aims to convert a source utterance into the voice of an unseen target speaker while preserving its linguistic content. Although recent systems have…
LPM 1.0: Video-based Character Performance Model
Ailing Zeng, Casper Yang, Chauncey Ge +22
Performance, the externalization of intent, emotion, and personality through visual, vocal, and temporal behavior, is what makes a character alive. Learning such performance from v…
TS-MLLM: A Multi-Modal Large Language Model-based Framework for Industrial Time-Series Big Data Analysis
Haiteng Wang, Yikang Li, Yunfei Zhu +3
Accurate analysis of industrial time-series big data is critical for the Prognostics and Health Management (PHM) of industrial equipment. While recent advancements in Large Languag…
VIP: Video Inpainting Pipeline for Real World Human Removal
Huiming Sun, Yikang Li, Kangning Yang +11
Inpainting for real-world human and pedestrian removal in high-resolution video clips presents significant challenges, particularly in achieving high-quality outcomes, ensuring tem…
VideoXum: Cross-modal Visual and Textural Summarization of Videos
Jingyang Lin, Hang Hua, Ming Chen +4
Video summarization aims to distill the most important information from a source video to produce either an abridged clip or a textual narrative. Traditionally, different methods h…