1 paper
Shaofeng Yin, Jiaxin Ge, Zora Zhiruo Wang +6
Vision-as-inverse-graphics, the concept of reconstructing images into editable programs, remains challenging for Vision-Language Models (VLMs), which inherently lack fine-grained s…