activity
20192022
most citedSimilar Scenes arouse Similar Emotions: Parallel Data Augmentation for Stylized Image Captioning

19 citations · 37 across the 10 of their papers we have counts for

collaborators

14 papers

cs.CL20226 cited

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

Qiming Peng, Yinxu Pan, Wenjin Wang +12

Recent years have witnessed the rise and success of pre-training techniques in visually-rich document understanding. However, most existing methods lack the systematic mining and u…

cs.CL2022

DAMO-NLP at NLPCC-2022 Task 2: Knowledge Enhanced Robust NER for Speech Entity Linking

Shen Huang, Yuchen Zhai, Xinwei Long +4

Speech Entity Linking aims to recognize and disambiguate named entities in spoken languages. Conventional methods suffer gravely from the unfettered speech styles and the noisy tra…

cs.CV20223 cited

ERNIE-mmLayout: Multi-grained MultiModal Transformer for Document Understanding

Wenjin Wang, Zhengjie Huang, Bin Luo +8

Recent efforts of multimodal Transformers have improved Visually Rich Document Understanding (VrDU) tasks via incorporating visual and textual information. However, existing approa…

cs.CV2022

Diverse Instance Discovery: Vision-Transformer for Instance-Aware Multi-Label Image Recognition

Yunqing Hu, Xuan Jin, Yin Zhang +5

Previous works on multi-label image recognition (MLIR) usually use CNNs as a starting point for research. In this paper, we take pure Vision Transformer (ViT) as the research base…

cs.CV202119 cited

Similar Scenes arouse Similar Emotions: Parallel Data Augmentation for Stylized Image Captioning

Guodun Li, Yuchen Zhai, Zehao Lin +1

Stylized image captioning systems aim to generate a caption not only semantically related to a given image but also consistent with a given style description. One of the biggest ch…

cs.CV2021

DRDF: Determining the Importance of Different Multimodal Information with Dual-Router Dynamic Framework

Haiwen Hong, Xuan Jin, Yin Zhang +4

In multimodal tasks, we find that the importance of text and image modal information is different for different input cases, and for this motivation, we propose a high-performance…