1 paper
Abdelrahman Abdallah, Mahmoud Abdalla, Mohammed Ali +1
Late-interaction vision-language retrievers represent each document page as many visual token embeddings and score queries with MaxSim. In systems such as ColPali, ColQwen, ColNomi…