1 paper
Jaeseok Byun, Taebaek Hwang, Jianlong Fu +1
Most of the currently existing vision and language pre-training (VLP) methods have mainly focused on how to extract and align vision and text features. In contrast to the mainstrea…