116 citations · 235 across the 15 of their papers we have counts for
4 papers · 1 filter
VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
Linjie Li, Jie Lei, Zhe Gan +12
Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily…
Gaze Perception in Humans and CNN-Based Model
Nicole X. Han, William Yang Wang, Miguel P. Eckstein
Making accurate inferences about other individuals' locus of attention is essential for human social interactions and will be important for AI to effectively interact with humans.…
Meta Module Network for Compositional Visual Reasoning
Wenhu Chen, Zhe Gan, Linjie Li +3
Neural Module Network (NMN) exhibits strong interpretability and compositionality thanks to its handcrafted neural modules with explicit multi-hop reasoning capability. However, mo…
REVERIE: Remote Embodied Visual Referring Expression in Real Indoor Environments
Yuankai Qi, Qi Wu, Peter Anderson +4
One of the long-term challenges of robotics is to enable robots to interact with humans in the visual world via natural language, as humans are visual animals that communicate thro…