2 papers
cs.LG2023
Cross-modal Prompts: Adapting Large Pre-trained Models for Audio-Visual Downstream Tasks
Haoyi Duan, Yan Xia, Mingze Zhou +3
In recent years, the deployment of large-scale pre-trained models in audio-visual downstream tasks has yielded remarkable outcomes. However, these models, primarily trained on sing…
cs.IR2023
Beyond Two-Tower Matching: Learning Sparse Retrievable Cross-Interactions for Recommendation
Liangcai Su, Fan Yan, Jieming Zhu +5
Two-tower models are a prevalent matching framework for recommendation, which have been widely deployed in industrial applications. The success of two-tower matching attributes to…