9 papers
K-EXAONE 2.0 Technical Report
Eunbi Choi, Kibong Choi, Sehyun Chun +74
This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward global frontier-scale foundatio…
Video-Oasis: Rethinking Evaluation of Video Understanding
Geuntaek Lim, Sungjune Park, Jaeyun Lee +5
The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performance stems from visual perception, linguistic reasoning, or knowle…
Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues
Beomchan Park, Seongho Kim, Hyunjun Kim +2
While Multimodal Large Language Models (MLLMs) have enhanced grounding capabilities in general scenes, their robustness in crowded scenes remains underexplored. Crowded scenes enta…
Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
Jeong Hun Yeo, Hyeongseop Rha, Sungjune Park +2
Audio is the primary modality for human communication and has driven the success of Automatic Speech Recognition (ASR) technologies. However, such audio-centric systems inherently…
Robust Egocentric Visual Attention Prediction Through Language-guided Scene Context-aware Learning
Sungjune Park, Hongda Mao, Qingshuang Chen +2
As the demand for analyzing egocentric videos grows, egocentric visual attention prediction, anticipating where a camera wearer will attend, has garnered increasing attention. Howe…
GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
Jeong Hun Yeo, Sangyun Chung, Sungjune Park +3
Long-video understanding remains a significant challenge for Multimodal Large Language Models (MLLMs) due to inherent token limitations and the complexity of capturing long-term te…