6 papers
Low-rank Prompt Interaction for Continual Vision-Language Retrieval
Weicai Yan, Ye Wang, Wang Lin +3
Research on continual learning in multi-modal tasks has been receiving increasing attention. However, most existing work overlooks the explicit cross-modal and cross-task interacti…
Causal Distillation for Alleviating Performance Heterogeneity in Recommender Systems
Shengyu Zhang, Ziqi Jiang, Jiangchao Yao +7
Recommendation performance usually exhibits a long-tail distribution over users -- a small portion of head users enjoy much more accurate recommendation services than the others. W…
SkinGEN: an Explainable Dermatology Diagnosis-to-Generation Framework with Interactive Vision-Language Models
Bo Lin, Yingjing Xu, Xuanwen Bao +3
With the continuous advancement of vision language models (VLMs) technology, remarkable research achievements have emerged in the dermatology field, the fourth most prevalent human…
Beyond Two-Tower Matching: Learning Sparse Retrievable Cross-Interactions for Recommendation
Liangcai Su, Fan Yan, Jieming Zhu +5
Two-tower models are a prevalent matching framework for recommendation, which have been widely deployed in industrial applications. The success of two-tower matching attributes to…
DisCover: Disentangled Music Representation Learning for Cover Song Identification
Jiahao Xun, Shengyu Zhang, Yanting Yang +7
In the field of music information retrieval (MIR), cover song identification (CSI) is a challenging task that aims to identify cover versions of a query song from a massive collect…
Gloss Attention for Gloss-free Sign Language Translation
Aoxiong Yin, Tianyun Zhong, Li Tang +3
Most sign language translation (SLT) methods to date require the use of gloss annotations to provide additional supervision information, however, the acquisition of gloss is not ea…