158 citations · 485 across the 44 of their papers we have counts for
15 papers · 1 filter
MNER-QG: An End-to-End MRC framework for Multimodal Named Entity Recognition with Query Grounding
Meihuizi Jia, Lei Shen, Xin Shen +5
Multimodal named entity recognition (MNER) is a critical step in information extraction, which aims to detect entity spans and classify them to corresponding entity types given a s…
SE-GAN: Skeleton Enhanced GAN-based Model for Brush Handwriting Font Generation
Shaozu Yuan, Ruixue Liu, Meng Chen +3
Previous works on font generation mainly focus on the standard print fonts where character's shape is stable and strokes are clearly separated. There is rare research on brush hand…
Cross-modal Contrastive Distillation for Instructional Activity Anticipation
Zhengyuan Yang, Jingen Liu, Jing Huang +4
In this study, we aim to predict the plausible future action steps given an observation of the past and study the task of instructional activity anticipation. Unlike previous antic…
ViDA-MAN: Visual Dialog with Digital Humans
Tong Shen, Jiawei Zuo, Fan Shi +7
We demonstrate ViDA-MAN, a digital-human agent for multi-modal interaction, which offers realtime audio-visual responses to instant speech inquiries. Compared to traditional text o…
Object-driven Text-to-Image Synthesis via Adversarial Training
Wenbo Li, Pengchuan Zhang, Lei Zhang +4
In this paper, we propose Object-driven Attentive Generative Adversarial Newtorks (Obj-GANs) that allow object-centered text-to-image synthesis for complex scenes. Following the tw…
The Neural Painter: Multi-Turn Image Generation
Ryan Y. Benmalek, Claire Cardie, Serge Belongie +2
In this work we combine two research threads from Vision/ Graphics and Natural Language Processing to formulate an image generation task conditioned on attributes in a multi-turn s…