collaborators

10 papers

cs.RO2026

Advancing Multi-Robot Networks via MLLM-Driven Sensing, Communication, and Computation: A Comprehensive Survey

Hyun Jong Yang, Howon Lee, Kyuhong Shim +10

Imagine advanced humanoid robots, powered by multimodal large language models (MLLMs), coordinating missions across industries like warehouse logistics, manufacturing, and safety r…

cs.CV2026

Revealing Multi-View Hallucination in Large Vision-Language Models

Wooje Park, Insu Lee, Soohyun Kim +4

Large vision-language models (LVLMs) are increasingly being applied to multi-view image inputs captured from diverse viewpoints. However, despite this growing use, current LVLMs of…

cs.RO2026

Adaptive Capacity Allocation for Vision Language Action Fine-tuning

Donghoon Kim, Minji Bae, Unghui Nam +4

Vision language action models (VLAs) are increasingly used for Physical AI, but deploying a pre-trained VLA model to unseen environments, embodiments, or tasks still requires adapt…

eess.IV2025

InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding

Minsoo Kim, Kyuhong Shim, Jungwook Choi +1

Modern multimodal large language models (MLLMs) can reason over hour-long video, yet their key-value (KV) cache grows linearly with time-quickly exceeding the fixed memory of phone…

cs.CV2025

Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs

Insu Lee, Wooje Park, Jaeyun Jang +3

Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by…

cs.CV2025

Unlocking Transfer Learning for Open-World Few-Shot Recognition

Byeonggeun Kim, Juntae Lee, Kyuhong Shim +1

Few-Shot Open-Set Recognition (FSOSR) targets a critical real-world challenge, aiming to categorize inputs into known categories, termed closed-set classes, while identifying open-…