collaborators

5 papers

cs.CV2026

Context Blindness in DPO: Mitigating Object Hallucination in MLLMs via Context-Calibrated Preference Optimization

Byungoh Ko, Jinyoung Park, Jongha Kim +3

Multimodal large language models (MLLMs) have made rapid progress, yet they still exhibit object hallucination, generating plausible but incorrect descriptions that are inconsisten…

cs.CV2026

Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning

Seung Hyup Baek, Jimin Lee, Hyeongkeun Lee +1

Dense Video Captioning (DVC) is a challenging multimodal task that involves temporally localizing multiple events within a video and describing them with natural language. While qu…

cs.SD2025

MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning

Kyeongha Rho, Hyeongkeun Lee, Jae Won Cho +1

In this paper, we propose Mixture of Layer-Wise Tokens (MoLT), a parameter- and memory-efficient adaptation framework for audio-visual learning. The key idea of MoLT is to replace…

cs.RO2025

Neural Brain: A Neuroscience-inspired Framework for Embodied Agents

Jian Liu, Xiongtao Shi, Thai Duy Nguyen +13

The rapid evolution of artificial intelligence (AI) has shifted from static, data-driven models to dynamic systems capable of perceiving and interacting with real-world environment…

cs.CV2025

Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization

Ji Soo Lee, Byungoh Ko, Jaewon Cho +3

In text-video retrieval, auxiliary captions are often used to enhance video understanding, bridging the gap between the modalities. While recent advances in multi-modal large langu…