85 citations · 361 across the 21 of their papers we have counts for
1 paper · 1 filter
Jia Guo, Chen Zhu, Yilun Zhao +4
Multi-modal representation learning by pretraining has become an increasing interest due to its easy-to-use and potential benefit for various Visual-and-Language~(V-L) tasks. Howev…