activity
20202025
most citedDeVLBert: Learning Deconfounded Visio-Linguistic Representations

64 citations · 122 across the 17 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2024

Semantic Alignment for Multimodal Large Language Models

Tao Wu, Mengze Li, Jingyuan Chen +6

Research on Multi-modal Large Language Models (MLLMs) towards the multi-image cross-modal instruction has received increasing attention and made significant progress, particularly…

cs.CV20241 cited

MetaCoCo: A New Few-Shot Classification Benchmark with Spurious Correlation

Min Zhang, Haoxuan Li, Fei Wu +1

Out-of-distribution (OOD) problems in few-shot classification (FSC) occur when novel classes sampled from testing distributions differ from base classes drawn from training distrib…

cs.CV20224 cited

Domain Generalization via Contrastive Causal Learning

Qiaowei Miao, Junkun Yuan, Kun Kuang

Domain Generalization (DG) aims to learn a model that can generalize well to unseen target domains from a set of source domains. With the idea of invariant causal mechanism, a lot…

cs.CV202064 cited

DeVLBert: Learning Deconfounded Visio-Linguistic Representations

Shengyu Zhang, Tan Jiang, Tan Wang +6

In this paper, we propose to investigate the problem of out-of-domain visio-linguistic pretraining, where the pretraining data distribution differs from that of downstream data on…

cs.CV202025 cited

Poet: Product-oriented Video Captioner for E-commerce

Shengyu Zhang, Ziqi Tan, Jin Yu +6

In e-commerce, a growing number of user-generated videos are used for product promotion. How to generate video descriptions that narrate the user-preferred product characteristics…

cs.CV202024 cited

Comprehensive Information Integration Modeling Framework for Video Titling

Shengyu Zhang, Ziqi Tan, Jin Yu +6

In e-commerce, consumer-generated videos, which in general deliver consumers' individual preferences for the different aspects of certain products, are massive in volume. To recomm…