2 papers
cs.IR2024
Attribute-driven Disentangled Representation Learning for Multimodal Recommendation
Zhenyang Li, Fan Liu, Yinwei Wei +3
Recommendation algorithms forecast user preferences by correlating user and item representations derived from historical interaction patterns. In pursuit of enhanced performance, m…
cs.CV2024
Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCR
Zhenyang Li, Yangyang Guo, Kejie Wang +3
Visual Commonsense Reasoning (VCR) calls for explanatory reasoning behind question answering over visual scenes. To achieve this goal, a model is required to provide an acceptable…