1 paper
Kim Yu-Ji, Dahye Lee, Kim Jun-Seong +6
Vision-Language Models (VLMs) have demonstrated strong reasoning capabilities over images and videos, yet their application to embodied scene understanding often constrained by the…