7 papers
From Saliency to Discriminability: Rank-Preserving Visual Token Pruning for VLM Rerankers
Siyi Liu, Hanjun Yang, Chenchen Zhang +7
Large vision-language models used as listwise rerankers must jointly process visual tokens from tens of candidates per query, making token pruning essential for practical deploymen…
SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
Zirong Chen, Fuda Ye, Kuan Zhang +7
Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mo…
RePair: Turning Retrieval Failures into Counterfactual Hard Pairs
Siyi Liu, Xiaorong Zhu, Enjun Du +6
Vision-language retrieval with CLIP-style dual encoders achieves strong cross-modal performance, yet practical accuracy often hinges on localized semantic distinctions where top-ra…
EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
Enjun Du, Siyi Liu, Zirong Chen +8
Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet exis…
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
Enjun Du, Hange Zhou, Chenxu Du +4
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely r…
Mixture of Length and Pruning Experts for Knowledge Graphs Reasoning
Enjun Du, Siyi Liu, Yongqi Zhang
Knowledge Graph (KG) reasoning, which aims to infer new facts from structured knowledge repositories, plays a vital role in Natural Language Processing (NLP) systems. Its effective…