1 paper
Shuxin Yang, Xinhan Di
There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multi-modal models fail to provide sati…