1 citations · 1 across the 9 of their papers we have counts for
8 papers · 1 filter
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
Junho Kim, Hosu Lee, James M. Rehg +2
Recent progress in video large language models (Video-LLMs) has enabled strong offline reasoning over long and complex videos. However, real-world deployments increasingly require…
Unified Reinforcement and Imitation Learning for Vision-Language Models
Byung-Kwan Lee, Ryo Hachiuma, Yong Man Ro +2
Vision-Language Models (VLMs) have achieved remarkable progress, yet their large scale often renders them impractical for resource-constrained environments. This paper introduces U…
Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images
Sungjune Park, Hyunjun Kim, Beomchan Park +1
Despite recent advancements in computer vision research, object detection in aerial images still suffers from several challenges. One primary challenge to be mitigated is the prese…
DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
Sungjune Park, Hyunjun Kim, Junho Kim +2
MLLMs have demonstrated significant visual understanding capabilities, yet their fine-grained visual perception in complex real-world scenarios, such as densely crowded public area…
VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models
Byung-Kwan Lee, Ryo Hachiuma, Yu-Chiang Frank Wang +2
The recent surge in high-quality visual instruction tuning samples from closed-source vision-language models (VLMs) such as GPT-4V has accelerated the release of open-source VLMs a…
Look Every Frame All at Once: Video-Mamba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
Hosu Lee, Junho Kim, Hyunjun Kim +1
With the growing scale and complexity of video data, efficiently processing long video sequences poses significant challenges due to the quadratic increase in memory and computatio…