1 paper
Fenglin Liu, Xian Wu, Shen Ge +4
Vision-and-language (V-L) tasks require the system to understand both vision content and natural language, thus learning fine-grained joint representations of vision and language (…