2 papers
cs.RO2026
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
Jaewoo Park, Minyoung Lee, Sukmin Seo +11
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends into acting: dropping an MLLM directly into a drone's control…
cs.CV2024
CREPE: Coordinate-Aware End-to-End Document Parser
Yamato Okamoto, Youngmin Baek, Geewook Kim +5
In this study, we formulate an OCR-free sequence generation model for visual document understanding (VDU). Our model not only parses text from document images but also extracts the…