2 papers
cs.CV2021
All Tokens Matter: Token Labeling for Training Better Vision Transformers
Zihang Jiang, Qibin Hou, Li Yuan +5
In this paper, we present token labeling -- a new training objective for training high-performance vision transformers (ViTs). Different from the standard training objective of ViT…
cs.CV2018
Holistic Multi-modal Memory Network for Movie Question Answering
Anran Wang, Anh Tuan Luu, Chuan-Sheng Foo +3
Answering questions according to multi-modal context is a challenging problem as it requires a deep integration of different data sources. Existing approaches only employ partial i…