3 papers
cs.CL2026
Where Do Multilingual Vision-Language Encoders Fail on Low-Resource Languages?
Donghoon Han, SungHyun Moon, Aidyn Zhakatayev +2
Recent multilingual vision--language encoders cover hundreds of languages in a single model, yet on two state-of-the-art instances retrieval on low-resource languages (LRL; e.g. Sw…
cs.CL2024
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
Donghoon Han, Eunhwan Park, Gisang Lee +2
The rapid expansion of multimedia content has made accurately retrieving relevant videos from large collections increasingly challenging. Recent advancements in text-video retrieva…
cs.CV2021
The U-Net based GLOW for Optical-Flow-free Video Interframe Generation
Saem Park, Donghoon Han, Nojun Kwak
Video frame interpolation is the task of creating an interframe between two adjacent frames along the time axis. So, instead of simply averaging two adjacent frames to create an in…