1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Mustafa Shukor, Dana Aubakirova, Francesco Capuano +11
Vision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics. Rather than t…