1 citations · 1 across the 2 of their papers we have counts for
4 papers
Beyond Bag-of-Patches: Learning Global Layout via Textual Supervision for Late-Interaction Visual Document Retrieval
Pascal Tilli, Mohsen Mesgar
Visual Document Retrieval (VDR) models mostly rely on late interaction architectures, in which documents are represented by a set of local patch embeddings and then matched against…
HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
Esra Dönmez, Pascal Tilli, Hsiu-Yu Yang +2
Image-Text-Matching (ITM) is one of the defacto methods of learning generalized representations from a large corpus in Vision and Language (VL). However, due to the weak associatio…
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
Lucas Möller, Pascal Tilli, Ngoc Thang Vu +1
Dual encoder architectures like Clip models map two types of inputs into a shared embedding space and predict similarities between them. Despite their wide application, it is, howe…
Discrete Subgraph Sampling for Interpretable Graph based Visual Question Answering
Pascal Tilli, Ngoc Thang Vu
Explainable artificial intelligence (XAI) aims to make machine learning models more transparent. While many approaches focus on generating explanations post-hoc, interpretable appr…