1 citations · 1 across the 2 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.CV2026
Beyond Bag-of-Patches: Learning Global Layout via Textual Supervision for Late-Interaction Visual Document Retrieval
Pascal Tilli, Mohsen Mesgar
Visual Document Retrieval (VDR) models mostly rely on late interaction architectures, in which documents are represented by a set of local patch embeddings and then matched against…
cs.CL2026★ 1 cited
HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
Esra Dönmez, Pascal Tilli, Hsiu-Yu Yang +2
Image-Text-Matching (ITM) is one of the defacto methods of learning generalized representations from a large corpus in Vision and Language (VL). However, due to the weak associatio…