3 papers
cs.CV2023
RTQ: Rethinking Video-language Understanding Based on Image-text Model
Xiao Wang, Yaoyu Li, Tian Gan +3
Recent advancements in video-language understanding have been established on the foundation of image-text models, resulting in promising outcomes due to the shared knowledge betwee…
cs.CV2023
Mutual Query Network for Multi-Modal Product Image Segmentation
Yun Guo, Wei Feng, Zheng Zhang +6
Product image segmentation is vital in e-commerce. Most existing methods extract the product image foreground only based on the visual modality, making it difficult to distinguish…
cs.CV2023
Relation-Aware Diffusion Model for Controllable Poster Layout Generation
Fengheng Li, An Liu, Wei Feng +8
Poster layout is a crucial aspect of poster design. Prior methods primarily focus on the correlation between visual content and graphic elements. However, a pleasant layout should…