3 papers
cs.CV2026
VINO: Video-driven Invariance for Non-contextual Objects via Structural Prior Guided De-contextualization
Seul-Ki Yeom, Marcel Simon, Eunbin Lee +1
Self-supervised learning (SSL) has made rapid progress, yet learned features often over-rely on contextual shortcuts-background textures and co-occurrence statistics. While video p…
cs.CV2025
Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception
Marcel Simon, Tae-Ho Kim, Seul-Ki Yeom
Self-supervised image encoders such as DINO have recently gained significant interest for learning robust visual features without labels. However, most SSL methods train on static…
cs.CV2024
UniForm: A Reuse Attention Mechanism Optimized for Efficient Vision Transformers on Edge Devices
Seul-Ki Yeom, Tae-Ho Kim
Transformer-based architectures have demonstrated remarkable success across various domains, but their deployment on edge devices remains challenging due to high memory and computa…