6 papers
Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval
Jihyun Lee, Cheol-Ho Cho, Woojin Jun +2
Zero-shot video moment retrieval aims to overcome the limitations of traditional approaches that require large-scale datasets annotated with text and its relevant temporal spans. D…
Analyzing the Training Dynamics of Image Restoration Transformers: A Revisit to Layer Normalization
MinKyu Lee, Sangeek Hyun, Woojin Jun +3
This work analyzes the training dynamics of Image Restoration (IR) Transformers and uncovers a critical yet overlooked issue: conventional LayerNorm (LN) drives feature magnitudes…
Mitigating Semantic Collapse in Partially Relevant Video Retrieval
WonJun Moon, MinSeok Jung, Gilhan Park +4
Partially Relevant Video Retrieval (PRVR) seeks videos where only part of the content matches a text query. Existing methods treat every annotated text-video pair as a positive and…
Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval
CH Cho, WJ Moon, W Jun +2
Partially Relevant Video Retrieval~(PRVR) aims to retrieve a video where a specific segment is relevant to a given text query. Typical training processes of PRVR assume a one-to-on…
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
WonJun Moon, Cheol-Ho Cho, Woojin Jun +5
In a retrieval system, simultaneously achieving search accuracy and efficiency is inherently challenging. This challenge is particularly pronounced in partially relevant video retr…
Auto-Encoded Supervision for Perceptual Image Super-Resolution
MinKyu Lee, Sangeek Hyun, Woojin Jun +1
This work tackles the fidelity objective in the perceptual super-resolution~(SR). Specifically, we address the shortcomings of pixel-level loss ($\mathcal{L}_\text{pix…