60 citations · 61 across the 3 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2022
DiMBERT: Learning Vision-Language Grounded Representations with Disentangled Multimodal-Attention
Fenglin Liu, Xian Wu, Shen Ge +4
Vision-and-language (V-L) tasks require the system to understand both vision content and natural language, thus learning fine-grained joint representations of vision and language (…
cs.CV2016★ 1 cited
An Automated CNN Recommendation System for Image Classification Tasks
Song Wang, Li Sun, Wei Fan +6
Nowadays the CNN is widely used in practical applications for image classification task. However the design of the CNN model is very professional work and which is very difficult f…