2 papers
cs.MM2024
Intelligent Text-Conditioned Music Generation
Zhouyao Xie, Nikhil Yadala, Xinyi Chen +1
CLIP (Contrastive Language-Image Pre-Training) is a multimodal neural network trained on (text, image) pairs to predict the most relevant text caption given an image. It has been u…
cs.IR2024
Making Recommender Systems More Knowledgeable: A Framework to Incorporate Side Information
Yukun Jiang, Leo Guo, Xinyi Chen +1
Session-based recommender systems typically focus on using only the triplet (user_id, timestamp, item_id) to make predictions of users' next actions. In this paper, we aim to utili…