47 citations · 120 across the 21 of their papers we have counts for
35 papers
LiveSeg: Unsupervised Multimodal Temporal Segmentation of Long Livestream Videos
Jielin Qiu, Franck Dernoncourt, Trung Bui +3
Livestream videos have become a significant part of online learning, where design, digital marketing, creative painting, and other skills are taught by experienced experts in the s…
Semantics-Consistent Cross-domain Summarization via Optimal Transport Alignment
Jielin Qiu, Jiacheng Zhu, Mengdi Xu +6
Multimedia summarization with multimodal output (MSMO) is a recently explored application in language grounding. It plays an essential role in real-world applications, i.e., automa…
MHMS: Multimodal Hierarchical Multimedia Summarization
Jielin Qiu, Jiacheng Zhu, Mengdi Xu +6
Multimedia summarization with multimodal output can play an essential role in real-world applications, i.e., automatically generating cover images and titles for news articles or p…
StyleBabel: Artistic Style Tagging and Captioning
Dan Ruta, Andrew Gilbert, Pranav Aggarwal +9
We present StyleBabel, a unique open access dataset of natural language captions and free-form tags describing the artistic style of over 135K digital artworks, collected via a nov…
StreamHover: Livestream Transcript Summarization and Annotation
Sangwoo Cho, Franck Dernoncourt, Tim Ganter +7
With the explosive growth of livestream broadcasting, there is an urgent need for new summarization technology that enables us to create a preview of streamed content and tap into…
Font Completion and Manipulation by Cycling Between Multi-Modality Representations
Ye Yuan, Wuyang Chen, Zhaowen Wang +4
Generating font glyphs of consistent style from one or a few reference glyphs, i.e., font completion, is an important task in topographical design. As the problem is more well-defi…