17 citations · 17 across the 2 of their papers we have counts for
2 papers
cs.CV2024
Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCR
Zhenyang Li, Yangyang Guo, Kejie Wang +3
Visual Commonsense Reasoning (VCR) calls for explanatory reasoning behind question answering over visual scenes. To achieve this goal, a model is required to provide an acceptable…
cs.IR2023★ 17 cited
Attribute-driven Disentangled Representation Learning for Multimodal Recommendation
Zhenyang Li, Fan Liu, Yinwei Wei +3
Recommendation algorithms forecast user preferences by correlating user and item representations derived from historical interaction patterns. In pursuit of enhanced performance, m…