1 citations · 1 across the 2 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Beyond Bag-of-Patches: Learning Global Layout via Textual Supervision for Late-Interaction Visual Document Retrieval
Pascal Tilli, Mohsen Mesgar
Visual Document Retrieval (VDR) models mostly rely on late interaction architectures, in which documents are represented by a set of local patch embeddings and then matched against…
cs.CV2025
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
Lucas Möller, Pascal Tilli, Ngoc Thang Vu +1
Dual encoder architectures like Clip models map two types of inputs into a shared embedding space and predict similarities between them. Despite their wide application, it is, howe…