38 citations · 38 across the 2 of their papers we have counts for
3 papers
cs.CV2021★ 38 cited
VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
Linjie Li, Jie Lei, Zhe Gan +12
Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily…
cs.CL2019
Hierarchical Graph Network for Multi-hop Question Answering
Yuwei Fang, Siqi Sun, Zhe Gan +3
In this paper, we present Hierarchical Graph Network (HGN) for multi-hop question answering. To aggregate clues from scattered texts across multiple paragraphs, a hierarchical grap…
cs.CV2018
An Empirical Analysis of Deep Audio-Visual Models for Speech Recognition
Devesh Walawalkar, Yihui He, Rohit Pillai
In this project, we worked on speech recognition, specifically predicting individual words based on both the video frames and audio. Empowered by convolutional neural networks, the…