4 citations · 6 across the 28 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification
Feixiang Liu, Likun Wang, Qiang Qiu +3
Spatial relation questions require a model to identify the queried subject and object before comparing their layout. Yet a VLM can recognize both entities and still answer from the…
cs.CV2026
Visual Credit Audit for Multimodal Spatial Reasoning
Feixiang Liu, Qiang Qiu, Lanbo Sun +3
Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit…
cs.CV2024
Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models
Shicheng Xu, Liang Pang, Yunchang Zhu +2
Vision-language alignment in Large Vision-Language Models (LVLMs) successfully enables LLMs to understand visual input. However, we find that existing vision-language alignment met…