2 papers
cs.CV2026
OrdinalBench: A Benchmark Dataset for Diagnosing Generalization Limits in Ordinal Number Understanding of Vision-Language Models
Yusuke Tozaki, Hisashi Miyamori
Vision-Language Models (VLMs) have advanced across multimodal benchmarks but still show clear gaps in ordinal number understanding, i.e., the ability to track relative positions an…
cs.CV2025
Can Visual Encoder Learn to See Arrows?
Naoyuki Terashita, Yusuke Tozaki, Hideaki Omote +4
The diagram is a visual representation of a relationship illustrated with edges (lines or arrows), which is widely used in industrial and scientific communication. Although recogni…