From the 1 of 32 linked papers with an AI index.
1 citations · 1 across the 14 of their papers we have counts for
3 papers · 1 filter
SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification
Feixiang Liu, Likun Wang, Qiang Qiu +3
Spatial relation questions require a model to identify the queried subject and object before comparing their layout. Yet a VLM can recognize both entities and still answer from the…
Visual Credit Audit for Multimodal Spatial Reasoning
Feixiang Liu, Qiang Qiu, Lanbo Sun +3
The paper introduces Visual Credit Audit (VCA), a method to quantify how much an image actually contributes to a multimodal model’s answer on spatial reasoning tasks, separating co…
Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models
Shicheng Xu, Liang Pang, Yunchang Zhu +2
Vision-language alignment in Large Vision-Language Models (LVLMs) successfully enables LLMs to understand visual input. However, we find that existing vision-language alignment met…