57 citations · 99 across the 7 of their papers we have counts for
13 papers
Cross-modal Contrastive Distillation for Instructional Activity Anticipation
Zhengyuan Yang, Jingen Liu, Jing Huang +4
In this study, we aim to predict the plausible future action steps given an observation of the past and study the task of instructional activity anticipation. Unlike previous antic…
TAP: Text-Aware Pre-training for Text-VQA and Text-Caption
Zhengyuan Yang, Yijuan Lu, Jianfeng Wang +6
In this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answ…
Pose-based Body Language Recognition for Emotion and Psychiatric Symptom Interpretation
Zhengyuan Yang, Amanda Kay, Yuncheng Li +2
Inspired by the human ability to infer emotions from body language, we propose an automated framework for body language based emotion recognition starting from regular RGB videos.…
Dynamic Context-guided Capsule Network for Multimodal Machine Translation
Huan Lin, Fandong Meng, Jinsong Su +5
Multimodal machine translation (MMT), which mainly focuses on enhancing text-only translation with visual features, has attracted considerable attention from both computer vision a…
Improving One-stage Visual Grounding by Recursive Sub-query Construction
Zhengyuan Yang, Tianlang Chen, Liwei Wang +1
We improve one-stage visual grounding by addressing current limitations on grounding long and complex queries. Existing one-stage methods encode the entire language query as a sing…
A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine Translation
Yongjing Yin, Fandong Meng, Jinsong Su +4
Multi-modal neural machine translation (NMT) aims to translate source sentences into a target language paired with images. However, dominant multi-modal NMT models do not fully exp…