1 paper
Dayong Liang, Changmeng Zheng, Zhiyuan Wen +3
Traditional scene graphs primarily focus on spatial relationships, limiting vision-language models' (VLMs) ability to reason about complex interactions in visual scenes. This paper…