activity
20242026
most citedSimple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval

30 citations · 30 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CL2026

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

Xiaolin Chen, Xuemeng Song, Wenhao Shi +3

Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task th…

cs.CV2026

FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning

Haokun Wen, Xuemeng Song, Xinghao Xie +3

Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is highly desired in practice.…

cs.CL2025

Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems

Xiaolin Chen, Xuemeng Song, Haokun Wen +3

Textual response generation is pivotal for multimodal \mbox{task-oriented} dialog systems, which aims to generate proper textual responses based on the multimodal context. While ex…

cs.CV2024

Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCR

Zhenyang Li, Yangyang Guo, Kejie Wang +3

Visual Commonsense Reasoning (VCR) calls for explanatory reasoning behind question answering over visual scenes. To achieve this goal, a model is required to provide an acceptable…

cs.MM202430 cited

Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval

Haokun Wen, Xuemeng Song, Xiaolin Chen +3

Composed image retrieval (CIR) aims to retrieve the target image based on a multimodal query, i.e., a reference image paired with corresponding modification text. Recent CIR studie…