4 papers · 1 filter
Harmonized Tabular-Image Fusion via Gradient-Aligned Alternating Learning
Longfei Huang, Yang Yang
Multimodal tabular-image fusion is an emerging task that has received increasing attention in various domains. However, existing methods may be hindered by gradient conflicts betwe…
Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge
Xiangyu Wu, Zhouyang Chi, Yang Yang +1
In this paper, we present our solution for the WSDM2023 Toloka Visual Question Answering Challenge. Inspired by the application of multimodal pre-trained models to various downstre…
TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt
Xiangyu Wu, Qing-Yuan Jiang, Yang Yang +3
The recent introduction of prompt tuning based on pre-trained vision-language models has dramatically improved the performance of multi-label image classification. However, some ex…
Learning to Rebalance Multi-Modal Optimization by Adaptively Masking Subnetworks
Yang Yang, Hongpeng Pan, Qing-Yuan Jiang +2
Multi-modal learning aims to enhance performance by unifying models from various modalities but often faces the "modality imbalance" problem in real data, leading to a bias towards…