1 citations · 1 across the 1 of their papers we have counts for
1 paper
Wenmo Qiu, Xinhan Di
There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multimodal models fail to provide satis…