1 citations · 1 across the 3 of their papers we have counts for
4 papers
Harmonized Tabular-Image Fusion via Gradient-Aligned Alternating Learning
Longfei Huang, Yang Yang
Multimodal tabular-image fusion is an emerging task that has received increasing attention in various domains. However, existing methods may be hindered by gradient conflicts betwe…
Multimodal Classification via Modal-Aware Interactive Enhancement
Qing-Yuan Jiang, Zhouyang Chi, Yang Yang
Due to the notorious modality imbalance problem, multimodal learning (MML) leads to the phenomenon of optimization imbalance, thus struggling to achieve satisfactory performance. R…
Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge
Xiangyu Wu, Zhouyang Chi, Yang Yang +1
In this paper, we present our solution for the WSDM2023 Toloka Visual Question Answering Challenge. Inspired by the application of multimodal pre-trained models to various downstre…
TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt
Xiangyu Wu, Qing-Yuan Jiang, Yang Yang +3
The recent introduction of prompt tuning based on pre-trained vision-language models has dramatically improved the performance of multi-label image classification. However, some ex…