1 paper
Jueqing Lu, Yuanyuan Qi, Xiaohao Yang +8
The performance of Visio-Language Transformers drops sharply when an input modality (e.g., image) is missing, because the model is forced to make predictions using incomplete infor…