2 papers
cs.CV2024
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
Chull Hwan Song, Taebaek Hwang, Jooyoung Yoon +2
Vision-language models (VLMs) have made significant strides in cross-modal understanding through large-scale paired datasets. However, in fashion domain, datasets often exhibit a d…
cs.CV2023
Conditional Cross Attention Network for Multi-Space Embedding without Entanglement in Only a SINGLE Network
Chull Hwan Song, Taebaek Hwang, Jooyoung Yoon +2
Many studies in vision tasks have aimed to create effective embedding spaces for single-label object prediction within an image. However, in reality, most objects possess multiple…