26 citations · 125 across the 21 of their papers we have counts for
25 papers
Language Conditioned Spatial Relation Reasoning for 3D Object Grounding
Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi +2
Localizing objects in 3D scenes based on natural language requires understanding and reasoning about spatial relations. In particular, it is often crucial to distinguish similar ob…
Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation
Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi +2
Following language instructions to navigate in unseen environments is a challenging problem for autonomous embodied agents. The agent not only needs to ground languages in visual s…
Product-oriented Machine Translation with Cross-modal Cross-lingual Pre-training
Yuqing Song, Shizhe Chen, Qin Jin +3
Translating e-commercial product descriptions, a.k.a product-oriented machine translation (PMT), is essential to serve e-shoppers all over the world. However, due to the domain spe…
Airbert: In-domain Pretraining for Vision-and-Language Navigation
Pierre-Louis Guhur, Makarand Tapaswi, Shizhe Chen +2
Vision-and-language navigation (VLN) aims to enable embodied agents to navigate in realistic environments using natural language instructions. Given the scarcity of domain-specific…
Elaborative Rehearsal for Zero-shot Action Recognition
Shizhe Chen, Dong Huang
The growing number of action classes has posed a new challenge for video understanding, making Zero-Shot Action Recognition (ZSAR) a thriving direction. The ZSAR task aims to recog…
Question-controlled Text-aware Image Captioning
Anwen Hu, Shizhe Chen, Qin Jin
For an image with multiple scene texts, different people may be interested in different text information. Current text-aware image captioning models are not able to generate distin…