416 citations · 843 across the 53 of their papers we have counts for
11 papers · 1 filter
LayoutNUWA: Revealing the Hidden Layout Expertise of Large Language Models
Zecheng Tang, Chenfei Wu, Juntao Li +1
Graphic layout generation, a growing research field, plays a significant role in user engagement and information perception. Existing methods primarily treat layout generation as a…
NÜWA-LIP: Language Guided Image Inpainting with Defect-free VQGAN
Minheng Ni, Chenfei Wu, Haoyang Huang +3
Language guided image inpainting aims to fill in the defective regions of an image under the guidance of text while keeping non-defective regions unchanged. However, the encoding p…
Hybrid Reasoning Network for Video-based Commonsense Captioning
Weijiang Yu, Jian Liang, Lei Ji +4
The task of video-based commonsense captioning aims to generate event-wise captions and meanwhile provide multiple commonsense descriptions (e.g., attribute, effect and intention)…
CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
Huaishao Luo, Lei Ji, Ming Zhong +4
Video-text retrieval plays an essential role in multi-modal research and has been widely used in many real-world web applications. The CLIP (Contrastive Language-Image Pre-training…
GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
Chenfei Wu, Lun Huang, Qianxi Zhang +5
Generating videos from text is a challenging task due to its high computational requirements for training and infinite possible answers for evaluation. Existing works typically exp…
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
Huaishao Luo, Lei Ji, Botian Shi +6
With the recent success of the pre-training technique for NLP and image-linguistic tasks, some video-linguistic pre-training works are gradually developed to improve video-text rel…