1 paper
Peter El Hachem, Ahmed Nassar, A. Said Gurbuz +2
Vision-Language Models (VLMs) parse documents end-to-end but frequently break down on layouts unlike those seen in training. We attribute this to a two-hop bottleneck: before the d…