15 papers · 1 filter
Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues
Beomchan Park, Seongho Kim, Hyunjun Kim +2
While Multimodal Large Language Models (MLLMs) have enhanced grounding capabilities in general scenes, their robustness in crowded scenes remains underexplored. Crowded scenes enta…
Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
Jeong Hun Yeo, Hyeongseop Rha, Sungjune Park +2
Audio is the primary modality for human communication and has driven the success of Automatic Speech Recognition (ASR) technologies. However, such audio-centric systems inherently…
Robust Egocentric Visual Attention Prediction Through Language-guided Scene Context-aware Learning
Sungjune Park, Hongda Mao, Qingshuang Chen +2
As the demand for analyzing egocentric videos grows, egocentric visual attention prediction, anticipating where a camera wearer will attend, has garnered increasing attention. Howe…
GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
Jeong Hun Yeo, Sangyun Chung, Sungjune Park +3
Long-video understanding remains a significant challenge for Multimodal Large Language Models (MLLMs) due to inherent token limitations and the complexity of capturing long-term te…
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
Jeong Hun Yeo, Minsu Kim, Chae Won Kim +2
We explore a novel zero-shot Audio-Visual Speech Recognition (AVSR) framework, dubbed Zero-AVSR, which enables speech recognition in target languages without requiring any audio-vi…
Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
Sungjune Park, Yeongyun Kim, Se Yeon Kim +1
Large Vision and Language Models (LVLMs) have shown strong performance across various vision-language tasks in natural image domains. However, their application to remote sensing (…