2 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CV2024
VidLA: Video-Language Alignment at Scale
Mamshad Nayeem Rizve, Fan Fei, Jayakrishnan Unnikrishnan +5
In this paper, we propose VidLA, an approach for video-language alignment at scale. There are two major limitations of previous video-language alignment approaches. First, they do…
cs.CL2023★ 1 cited
Graph-Aware Language Model Pre-Training on a Large Graph Corpus Can Help Multiple Graph Applications
Han Xie, Da Zheng, Jun Ma +9
Model pre-training on large text corpora has been demonstrated effective for various downstream applications in the NLP domain. In the graph mining domain, a similar analogy can be…
cs.LG2023★ 2 cited
Understanding and Constructing Latent Modality Structures in Multi-modal Representation Learning
Qian Jiang, Changyou Chen, Han Zhao +6
Contrastive loss has been increasingly used in learning representations from multiple modalities. In the limit, the nature of the contrastive loss encourages modalities to exactly…