12 citations · 24 across the 10 of their papers we have counts for
11 papers
Attract me to Buy: Advertisement Copywriting Generation with Multimodal Multi-structured Information
Zhipeng Zhang, Xinglin Hou, Kai Niu +5
Recently, online shopping has gradually become a common way of shopping for people all over the world. Wonderful merchandise advertisements often attract more people to buy. These…
Dual-Level Decoupled Transformer for Video Captioning
Yiqi Gao, Xinglin Hou, Wei Suo +4
Video captioning aims to understand the spatio-temporal semantic concept of the video and generate descriptive sentences. The de-facto approach to this task dictates a text generat…
CapOnImage: Context-driven Dense-Captioning on Image
Yiqi Gao, Xinglin Hou, Yuanmeng Zhang +3
Existing image captioning systems are dedicated to generating narrative captions for images, which are spatially detached from the image in presentation. However, texts can also be…
Self-Supervised Text Erasing with Controllable Image Synthesis
Gangwei Jiang, Shiyao Wang, Tiezheng Ge +3
Recent efforts on scene text erasing have shown promising results. However, existing methods require rich yet costly label annotations to obtain robust models, which limits the use…
Structure-Aware Motion Transfer with Deformable Anchor Model
Jiale Tao, Biao Wang, Borun Xu +4
Given a source image and a driving video depicting the same object type, the motion transfer task aims to generate a video by learning the motion from the driving video while prese…
Learning Pixel-Level Distinctions for Video Highlight Detection
Fanyue Wei, Biao Wang, Tiezheng Ge +3
The goal of video highlight detection is to select the most attractive segments from a long video to depict the most interesting parts of the video. Existing methods typically focu…