19 citations · 79 across the 25 of their papers we have counts for
8 papers · 1 filter
LiveSeg: Unsupervised Multimodal Temporal Segmentation of Long Livestream Videos
Jielin Qiu, Franck Dernoncourt, Trung Bui +3
Livestream videos have become a significant part of online learning, where design, digital marketing, creative painting, and other skills are taught by experienced experts in the s…
Semantics-Consistent Cross-domain Summarization via Optimal Transport Alignment
Jielin Qiu, Jiacheng Zhu, Mengdi Xu +6
Multimedia summarization with multimodal output (MSMO) is a recently explored application in language grounding. It plays an essential role in real-world applications, i.e., automa…
MHMS: Multimodal Hierarchical Multimedia Summarization
Jielin Qiu, Jiacheng Zhu, Mengdi Xu +6
Multimedia summarization with multimodal output can play an essential role in real-world applications, i.e., automatically generating cover images and titles for news articles or p…
Learning by Planning: Language-Guided Global Image Editing
Jing Shi, Ning Xu, Yihang Xu +3
Recently, language-guided global image editing draws increasing attention with growing application potentials. However, previous GAN-based methods are not only confined to domain-s…
A Benchmark and Baseline for Language-Driven Image Editing
Jing Shi, Ning Xu, Trung Bui +3
Language-driven image editing can significantly save the laborious image editing work and be friendly to the photography novice. However, most similar work can only deal with a spe…
PhraseCut: Language-based Image Segmentation in the Wild
Chenyun Wu, Zhe Lin, Scott Cohen +2
We consider the problem of segmenting image regions given a natural language phrase, and study it on a novel dataset of 77,262 images and 345,486 phrase-region pairs. Our dataset i…