benchmark evaluation 1image-text alignment 1large language models 1multimodal spatial reasoning 1visual credit audit 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification
Feixiang Liu, Likun Wang, Qiang Qiu +3
Spatial relation questions require a model to identify the queried subject and object before comparing their layout. Yet a VLM can recognize both entities and still answer from the…
cs.CV2026
Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference
Feixiang Liu, Qiang Qiu, Hao Zhang +1
Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct fail…
cs.CV2026
Visual Credit Audit for Multimodal Spatial Reasoning
Feixiang Liu, Qiang Qiu, Lanbo Sun +3
The paper introduces Visual Credit Audit (VCA), a method to quantify how much an image actually contributes to a multimodal model’s answer on spatial reasoning tasks, separating co…