13 citations · 15 across the 3 of their papers we have counts for
4 papers
Text2Layer: Layered Image Generation using Latent Diffusion Model
Xinyang Zhang, Wentian Zhao, Xin Lu +1
Layer compositing is one of the most popular image editing workflows among both amateurs and professionals. Motivated by the success of diffusion models, we explore layer compositi…
Boosting Entity-aware Image Captioning with Multi-modal Knowledge Graph
Wentian Zhao, Yao Hu, Heda Wang +2
Entity-aware image captioning aims to describe named entities and events related to the image by utilizing the background knowledge in the associated article. This task remains cha…
Adaptive Recursive Circle Framework for Fine-grained Action Recognition
Hanxi Lin, Xinxiao Wu, Jiebo Luo
How to model fine-grained spatial-temporal dynamics in videos has been a challenging problem for action recognition. It requires learning deep and rich features with superior disti…
Relational Reasoning using Prior Knowledge for Visual Captioning
Jingyi Hou, Xinxiao Wu, Yayun Qi +3
Exploiting relationships among objects has achieved remarkable progress in interpreting images or videos by natural language. Most existing methods resort to first detecting object…