1 paper
Haruto Yoshida, Keito Kudo, Yoichi Aoki +4
Large vision-language models (LVLMs) demonstrate strong performance on diagram understanding benchmarks, yet they still struggle with understanding relationships between elements,…