6 citations · 6 across the 1 of their papers we have counts for
1 paper · 1 filter
Philipp J. Rösch, Jindřich Libovický
In most Vision-Language models (VL), the understanding of the image structure is enabled by injecting the position information (PI) about objects in the image. In our case study of…