2 papers
cs.RO2026
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
Jaewoo Park, Minyoung Lee, Sukmin Seo +11
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends into acting: dropping an MLLM directly into a drone's control…
cs.CV2026
Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition
Sukmin Seo, Geewook Kim
Temporal grounding--returning the interval for a natural-language query over a video--is the language interface to long-form video, yet has been studied on short video…