7 papers · 1 filter
Are Multimodal Embeddings Truly Beneficial for Recommendation? A Deep Dive into Whole vs. Individual Modalities
Yu Ye, Junchen Fu, Yu Song +2
Multimodal recommendation has emerged as a mainstream paradigm, typically leveraging text and visual embeddings extracted from pre-trained models such as Sentence-BERT, Vision Tran…
Video-Bench: Human-Aligned Video Generation Benchmark
Hui Han, Siyuan Li, Jiaqi Chen +10
Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video g…
The 1st EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval
Junchen Fu, Xuri Ge, Xin Xin +5
Multimodal representation learning has garnered significant attention in the AI community, largely due to the success of large pre-trained multimodal foundation models like LLaMA,…
Teach Me How to Denoise: A Universal Framework for Denoising Multi-modal Recommender Systems via Guided Calibration
Hongji Li, Hanwen Du, Youhua Li +5
The surge in multimedia content has led to the development of Multi-Modal Recommender Systems (MMRecs), which use diverse modalities such as text, images, videos, and audio for mor…
Multimodal Representation Learning Techniques for Comprehensive Facial State Analysis
Kaiwen Zheng, Xuri Ge, Junchen Fu +2
Multimodal foundation models have significantly improved feature representation by integrating information from multiple modalities, making them highly suitable for a broader set o…
CROSSAN: Towards Efficient and Effective Adaptation of Multiple Multimodal Foundation Models for Sequential Recommendation
Junchen Fu, Yongxin Ni, Joemon M. Jose +4
In this paper, we explore a less-studied yet practically important problem: how to efficiently and effectively adapt multiple (2) multimodal foundation models (MFMs) for the seq…