1 paper
Haochen Huang, Jiahuan Pei, Yue Su +11
Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when fine-grained spatial reasoning and precis…