activity
20182022
most citedDynamic Context-guided Capsule Network for Multimodal Machine Translation

57 citations · 99 across the 7 of their papers we have counts for

collaborators

13 papers

cs.CV20221 cited

Cross-modal Contrastive Distillation for Instructional Activity Anticipation

Zhengyuan Yang, Jingen Liu, Jing Huang +4

In this study, we aim to predict the plausible future action steps given an observation of the past and study the task of instructional activity anticipation. Unlike previous antic…

cs.CV202019 cited

TAP: Text-Aware Pre-training for Text-VQA and Text-Caption

Zhengyuan Yang, Yijuan Lu, Jianfeng Wang +6

In this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answ…

cs.CV2020

Pose-based Body Language Recognition for Emotion and Psychiatric Symptom Interpretation

Zhengyuan Yang, Amanda Kay, Yuncheng Li +2

Inspired by the human ability to infer emotions from body language, we propose an automated framework for body language based emotion recognition starting from regular RGB videos.…

cs.CL202057 cited

Dynamic Context-guided Capsule Network for Multimodal Machine Translation

Huan Lin, Fandong Meng, Jinsong Su +5

Multimodal machine translation (MMT), which mainly focuses on enhancing text-only translation with visual features, has attracted considerable attention from both computer vision a…

cs.CV202017 cited

Improving One-stage Visual Grounding by Recursive Sub-query Construction

Zhengyuan Yang, Tianlang Chen, Liwei Wang +1

We improve one-stage visual grounding by addressing current limitations on grounding long and complex queries. Existing one-stage methods encode the entire language query as a sing…

cs.CL20205 cited

A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine Translation

Yongjing Yin, Fandong Meng, Jinsong Su +4

Multi-modal neural machine translation (NMT) aims to translate source sentences into a target language paired with images. However, dominant multi-modal NMT models do not fully exp…