2 citations · 2 across the 2 of their papers we have counts for
4 papers
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
Yongqi Li, Hongru Cai, Wenjie Wang +5
Text-to-image retrieval is a fundamental task in multimedia processing, aiming to retrieve semantically relevant cross-modal content. Traditional studies have typically approached…
Fine-tuning Multimodal Large Language Models for Product Bundling
Xiaohao Liu, Jie Wu, Zhulin Tao +3
Recent advances in product bundling have leveraged multimodal information through sophisticated encoders, but remain constrained by limited semantic understanding and a narrow scop…
Data-efficient Fine-tuning for LLM-based Recommendation
Xinyu Lin, Wenjie Wang, Yongqi Li +4
Leveraging Large Language Models (LLMs) for recommendation has recently garnered considerable attention, where fine-tuning plays a key role in LLMs' adaptation. However, the cost o…
Instilling Multi-round Thinking to Text-guided Image Generation
Lidong Zeng, Zhedong Zheng, Yinwei Wei +1
This paper delves into the text-guided image editing task, focusing on modifying a reference image according to user-specified textual feedback to embody specific attributes. Despi…