29 citations · 30 across the 4 of their papers we have counts for
6 papers
Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
Xiaolin Chen, Xuemeng Song, Haokun Wen +3
Textual response generation is pivotal for multimodal \mbox{task-oriented} dialog systems, which aims to generate proper textual responses based on the multimodal context. While ex…
OFFSET: Segmentation-based Focus Shift Revision for Composed Image Retrieval
Zhiwei Chen, Yupeng Hu, Zixu Li +3
Composed Image Retrieval (CIR) represents a novel retrieval paradigm that is capable of expressing users' intricate retrieval requirements flexibly. It enables the user to give a m…
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
Haoqiang Lin, Haokun Wen, Xuemeng Song +3
Composed Image Retrieval (CIR) allows users to search target images with a multimodal query, comprising a reference image and a modification text that describes the user's modifica…
A Comprehensive Survey on Composed Image Retrieval
Xuemeng Song, Haoqiang Lin, Haokun Wen +3
Composed Image Retrieval (CIR) is an emerging yet challenging task that allows users to search for target images using a multimodal query, comprising a reference image and a modifi…
Content-aware Balanced Spectrum Encoding in Masked Modeling for Time Series Classification
Yudong Han, Haocong Wang, Yupeng Hu +3
Due to the superior ability of global dependency, transformer and its variants have become the primary choice in Masked Time-series Modeling (MTM) towards time-series classificatio…
Vision-guided and Mask-enhanced Adaptive Denoising for Prompt-based Image Editing
Kejie Wang, Xuemeng Song, Meng Liu +2
Text-to-image diffusion models have demonstrated remarkable progress in synthesizing high-quality images from text prompts, which boosts researches on prompt-based image editing th…