1 citations · 1 across the 6 of their papers we have counts for
6 papers
Unveiling and Controlling Anomalous Attention Distribution in Transformers
Ruiqing Yan, Xingbo Du, Haoyu Deng +7
With the advent of large models based on the Transformer architecture, researchers have observed an anomalous phenomenon in the Attention mechanism--there is a very high attention…
Jaeger: A Concatenation-Based Multi-Transformer VQA Model
Jieting Long, Zewei Shi, Penghao Jiang +1
Document-based Visual Question Answering poses a challenging task between linguistic sense disambiguation and fine-grained multimodal retrieval. Although there has been encouraging…
Device Tuning for Multi-Task Large Model
Penghao Jiang, Xuanchen Hou, Yinsi Zhou
Unsupervised pre-training approaches have achieved great success in many fields such as Computer Vision (CV), Natural Language Processing (NLP) and so on. However, compared to typi…
Robust Meta Learning for Image based tasks
Penghao Jiang, Xin Ke, ZiFeng Wang +1
A machine learning model that generalizes well should obtain low errors on unseen test examples. Thus, if we learn an optimal model in training data, it could have better generaliz…
Invariant Meta Learning for Out-of-Distribution Generalization
Penghao Jiang, Ke Xin, Zifeng Wang +1
Modern deep learning techniques have illustrated their excellent capabilities in many areas, but relies on large training data. Optimization-based meta-learning train a model on a…
Deep Transfer Tensor Factorization for Multi-View Learning
Penghao Jiang, Ke Xin, Chunxi Li
This paper studies the data sparsity problem in multi-view learning. To solve data sparsity problem in multiview ratings, we propose a generic architecture of deep transfer tensor…