activity
20212026
most citedVision-Language Pre-Training with Triple Contrastive Learning

14 citations · 42 across the 18 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Zhiheng Liu, Weiming Ren, Xiaoke Huang +12

Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creating misalignment between the t…

cs.CV2024

Diffusion Models For Multi-Modal Generative Modeling

Changyou Chen, Han Ding, Bunyamin Sisman +5

Diffusion-based generative modeling has been achieving state-of-the-art results on various generation tasks. Most diffusion models, however, are limited to a single-generation mode…

cs.CV2024

VidLA: Video-Language Alignment at Scale

Mamshad Nayeem Rizve, Fan Fei, Jayakrishnan Unnikrishnan +5

In this paper, we propose VidLA, an approach for video-language alignment at scale. There are two major limitations of previous video-language alignment approaches. First, they do…

cs.CV2022★ 5 cited

Multi-modal Alignment using Representation Codebook

Jiali Duan, Liqun Chen, Son Tran +4

Aligning signals from different modalities is an important step in vision-language representation learning as it affects the performance of later stages such as cross-modality fusi…

cs.CV2022★ 14 cited

Vision-Language Pre-Training with Triple Contrastive Learning

Jinyu Yang, Jiali Duan, Son Tran +6

Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attrib…

cs.CV2021

MLIM: Vision-and-Language Model Pre-training with Masked Language and Image Modeling

Tarik Arici, Mehmet Saygin Seyfioglu, Tal Neiman +5

Vision-and-Language Pre-training (VLP) improves model performance for downstream tasks that require image and text inputs. Current VLP approaches differ on (i) model architecture (…