3 papers
cs.CV2026
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
Donghoon Han, Eunhwan Park, Seunghyeon Seo
Dense image retrieval is accurate but offers limited interpretability and attribution, and it can be compute-intensive at scale. We present \textbf{BM25-V}, which applies Okapi BM2…
cs.CL2024
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
Donghoon Han, Eunhwan Park, Gisang Lee +2
The rapid expansion of multimedia content has made accurately retrieving relevant videos from large collections increasingly challenging. Recent advancements in text-video retrieva…
cs.CV2024
Unleash the Potential of CLIP for Video Highlight Detection
Donghoon Han, Seunghyeon Seo, Eunhwan Park +2
Multimodal and large language models (LLMs) have revolutionized the utilization of open-world knowledge, unlocking novel potentials across various tasks and applications. Among the…