1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.RO2026
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
Jaewoo Park, Minyoung Lee, Sukmin Seo +11
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends into acting: dropping an MLLM directly into a drone's control…
cs.CL2024
SCANNER: Knowledge-Enhanced Approach for Robust Multi-modal Named Entity Recognition of Unseen Entities
Hyunjong Ok, Taeho Kil, Sukmin Seo +1
Recent advances in named entity recognition (NER) have pushed the boundary of the task to incorporate visual signals, leading to many variants, including multi-modal NER (MNER) or…
cs.CV2023★ 1 cited
Towards Unified Scene Text Spotting based on Sequence Generation
Taeho Kil, Seonghyeon Kim, Sukmin Seo +2
Sequence generation models have recently made significant progress in unifying various vision tasks. Although some auto-regressive models have demonstrated promising results in end…