1 paper · 1 filter
Xuejiao Tang, Wenbin Zhang
Given a question-image input, the Visual Commonsense Reasoning (VCR) model can predict an answer with the corresponding rationale, which requires inference ability from the real wo…