1 paper
Jiaye Feng, Qixiang Yin, Yuankun Liu +2
Scene Graph Generation (SGG) structures visual scenes as graphs of objects and their relations. While Multimodal Large Language Models (MLLMs) have advanced end-to-end SGG, current…