14 citations · 24 across the 5 of their papers we have counts for
5 papers
Multi-modal Alignment using Representation Codebook
Jiali Duan, Liqun Chen, Son Tran +4
Aligning signals from different modalities is an important step in vision-language representation learning as it affects the performance of later stages such as cross-modality fusi…
Vision-Language Pre-Training with Triple Contrastive Learning
Jinyu Yang, Jiali Duan, Son Tran +6
Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attrib…
Magic Pyramid: Accelerating Inference with Early Exiting and Token Pruning
Xuanli He, Iman Keivanloo, Yi Xu +4
Pre-training and then fine-tuning large language models is commonly used to achieve state-of-the-art performance in natural language processing (NLP) tasks. However, most pre-train…
MLIM: Vision-and-Language Model Pre-training with Masked Language and Image Modeling
Tarik Arici, Mehmet Saygin Seyfioglu, Tal Neiman +5
Vision-and-Language Pre-training (VLP) improves model performance for downstream tasks that require image and text inputs. Current VLP approaches differ on (i) model architecture (…
Tiering as a Stochastic Submodular Optimization Problem
Hyokun Yun, Michael Froh, Roshan Makhijani +3
Tiering is an essential technique for building large-scale information retrieval systems. While the selection of documents for high priority tiers critically impacts the efficiency…