1 paper · 1 filter
Jaewoo Park, Minyoung Lee, Sukmin Seo +11
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends into acting: dropping an MLLM directly into a drone's control…