2 papers
cs.CV2026
Structure-Token Evidence-Anchored Reasoning for Scientific Chart Understanding
Alberlucia Rafael Soarez, Camila Ferreira, Daniel Kim +2
Scientific charts encode quantities in axes, legends, and geometric marks, yet large vision-language models still treat them as natural photographs. Visual in-context examples do n…
eess.IV2026
SpaCE: Rethinking Spatial Capacity and Generalization in Multi-Frame Multimodal Large Language Models
Mariana Costa, Camila Ferreira, Alberlucia Rafael Soarez +1
Multi-modal large language models (MLLMs) have achieved remarkable empirical progress in spatial understanding through large-scale training on spatial visual question answering dat…