32 citations · 112 across the 14 of their papers we have counts for
28 papers
LiveSeg: Unsupervised Multimodal Temporal Segmentation of Long Livestream Videos
Jielin Qiu, Franck Dernoncourt, Trung Bui +3
Livestream videos have become a significant part of online learning, where design, digital marketing, creative painting, and other skills are taught by experienced experts in the s…
Semantics-Consistent Cross-domain Summarization via Optimal Transport Alignment
Jielin Qiu, Jiacheng Zhu, Mengdi Xu +6
Multimedia summarization with multimodal output (MSMO) is a recently explored application in language grounding. It plays an essential role in real-world applications, i.e., automa…
MHMS: Multimodal Hierarchical Multimedia Summarization
Jielin Qiu, Jiacheng Zhu, Mengdi Xu +6
Multimedia summarization with multimodal output can play an essential role in real-world applications, i.e., automatically generating cover images and titles for news articles or p…
Font Completion and Manipulation by Cycling Between Multi-Modality Representations
Ye Yuan, Wuyang Chen, Zhaowen Wang +4
Generating font glyphs of consistent style from one or a few reference glyphs, i.e., font completion, is an important task in topographical design. As the problem is more well-defi…
Black-Box Diagnosis and Calibration on GAN Intra-Mode Collapse: A Pilot Study
Zhenyu Wu, Zhaowen Wang, Ye Yuan +3
Generative adversarial networks (GANs) nowadays are capable of producing images of incredible realism. One concern raised is whether the state-of-the-art GAN's learned distribution…
Rethinking Text Segmentation: A Novel Dataset and A Text-Specific Refinement Approach
Xingqian Xu, Zhifei Zhang, Zhaowen Wang +3
Text segmentation is a prerequisite in many real-world text-related tasks, e.g., text style transfer, and scene text removal. However, facing the lack of high-quality datasets and…