4 papers
From Saliency to Discriminability: Rank-Preserving Visual Token Pruning for VLM Rerankers
Siyi Liu, Hanjun Yang, Chenchen Zhang +7
Large vision-language models used as listwise rerankers must jointly process visual tokens from tens of candidates per query, making token pruning essential for practical deploymen…
SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
Zirong Chen, Fuda Ye, Kuan Zhang +7
Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mo…
RePair: Turning Retrieval Failures into Counterfactual Hard Pairs
Siyi Liu, Xiaorong Zhu, Enjun Du +6
Vision-language retrieval with CLIP-style dual encoders achieves strong cross-modal performance, yet practical accuracy often hinges on localized semantic distinctions where top-ra…
EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
Enjun Du, Siyi Liu, Zirong Chen +8
Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet exis…