42 citations · 42 across the 2 of their papers we have counts for
1 paper · 1 filter
Haochen Huang, Yue Su, Xin Sun +11
Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when fine-grained spatial reasoning and precis…