2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Weixin Feng, Xingyuan Bu, Chenchen Zhang +1
Multimodal supervision has achieved promising results in many visual language understanding tasks, where the language plays an essential role as a hint or context for recognizing a…