99 citations · 127 across the 8 of their papers we have counts for
7 papers · 1 filter
Vision-Language Pre-Training with Triple Contrastive Learning
Jinyu Yang, Jiali Duan, Son Tran +6
Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attrib…
MLIM: Vision-and-Language Model Pre-training with Masked Language and Image Modeling
Tarik Arici, Mehmet Saygin Seyfioglu, Tal Neiman +5
Vision-and-Language Pre-training (VLP) improves model performance for downstream tasks that require image and text inputs. Current VLP approaches differ on (i) model architecture (…
SLADE: A Self-Training Framework For Distance Metric Learning
Jiali Duan, Yen-Liang Lin, Son Tran +2
Most existing distance metric learning approaches use fully labeled data to learn the sample similarities in an embedding space. We present a self-training framework, SLADE, to imp…
Deep Auto-Encoders with Sequential Learning for Multimodal Dimensional Emotion Recognition
Dung Nguyen, Duc Thanh Nguyen, Rui Zeng +5
Multimodal dimensional emotion recognition has drawn a great attention from the affective computing community and numerous schemes have been extensively investigated, making a sign…
Joint Deep Cross-Domain Transfer Learning for Emotion Recognition
Dung Nguyen, Sridha Sridharan, Duc Thanh Nguyen +4
Deep learning has been applied to achieve significant progress in emotion recognition. Despite such substantial progress, existing approaches are still hindered by insufficient tra…
Fashion Outfit Complementary Item Retrieval
Yen-Liang Lin, Son Tran, Larry S. Davis
Complementary fashion item recommendation is critical for fashion outfit completion. Existing methods mainly focus on outfit compatibility prediction but not in a retrieval setting…