7 papers
Stream-aware Side Adaptation for Large Pre-trained Multimodal Embedding Models in Sequential Recommendation
Junchen Fu, Kaiwen Zheng, Ioannis Arapakis +4
Recently, large pretrained multimodal embedding models such as Qwen3-VL Embedding have shown strong promise for sequential recommendation, as they provide reusable semantic item re…
MCoT-MVS: Multi-level Vision Selection by Multi-modal Chain-of-Thought Reasoning for Composed Image Retrieval
Xuri Ge, Chunhao Wang, Xindi Wang +3
Composed Image Retrieval (CIR) aims to retrieve target images based on a reference image and modified texts. However, existing methods often struggle to extract the correct semanti…
Identifying and Transferring Reasoning-Critical Neurons: Improving LLM Inference Reliability via Activation Steering
Fangan Dong, Zuming Yan, Xuri Ge +7
Despite the strong reasoning capabilities of recent large language models (LLMs), achieving reliable performance on challenging tasks often requires post-training or computationall…
LLMPopcorn: Exploring LLMs as Assistants for Popular Micro-video Generation
Junchen Fu, Xuri Ge, Kaiwen Zheng +5
In an era where micro-videos dominate platforms like TikTok and YouTube, AI-generated content is nearing cinematic quality. The next frontier is using large language models (LLMs)…
Efficient and Effective Adaptation of Multimodal Foundation Models in Sequential Recommendation
Junchen Fu, Xuri Ge, Xin Xin +5
Multimodal foundation models (MFMs) have revolutionized sequential recommender systems through advanced representation learning. While Parameter-efficient Fine-tuning (PEFT) is com…
The 1st EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval
Junchen Fu, Xuri Ge, Xin Xin +5
Multimodal representation learning has garnered significant attention in the AI community, largely due to the success of large pre-trained multimodal foundation models like LLaMA,…