1 paper
Hanyu Zhou, Gim Hee Lee
Vision-language models (VLMs) have demonstrated strong performance in 2D scene understanding and generation, but extending this unification to the physical world remains an open ch…