2 citations · 3 across the 29 of their papers we have counts for
1 paper · 1 filter
Wei Dai, Haoyu Wang, Honghao Chang +4
Vision Language Model (VLM) typically assume complete modality input during inference. However, their effectiveness drops sharply when certain modalities are unavailable or incompl…