32 citations · 93 across the 22 of their papers we have counts for
6 papers · 2 filters
COMPOSER: Compositional Reasoning of Group Activity in Videos with Keypoint-Only Modality
Honglu Zhou, Asim Kadav, Aviv Shamsian +6
Group Activity Recognition detects the activity collectively performed by a group of actors, which requires compositional reasoning of actors and objects. We approach the task by m…
Out-of-Domain Generalization from a Single Source: An Uncertainty Quantification Approach
Xi Peng, Fengchun Qiao, Long Zhao
We are concerned with a worst-case scenario in model generalization, in the sense that a model aims to perform well on many unseen domains while there is only one single domain ava…
Improved Transformer for High-Resolution GANs
Long Zhao, Zizhao Zhang, Ting Chen +2
Attention-based models, exemplified by the Transformer, can effectively model long range dependency, but suffer from the quadratic complexity of self-attention operation, making th…
Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual Understanding
Zizhao Zhang, Han Zhang, Long Zhao +3
Hierarchical structures are popular in recent vision transformers, however, they require sophisticated designs and massive datasets to work well. In this paper, we explore the idea…
More Than Just Attention: Improving Cross-Modal Attentions with Contrastive Constraints for Image-Text Matching
Yuxiao Chen, Jianbo Yuan, Long Zhao +4
Cross-modal attention mechanisms have been widely applied to the image-text matching task and have achieved remarkable improvements thanks to its capability of learning fine-graine…
SMIL: Multimodal Learning with Severely Missing Modality
Mengmeng Ma, Jian Ren, Long Zhao +3
A common assumption in multimodal learning is the completeness of training data, i.e., full modalities are available in all training examples. Although there exists research endeav…