7 citations · 13 across the 4 of their papers we have counts for
4 papers
A Case Study on Combining ASR and Visual Features for Generating Instructional Video Captions
Jack Hessel, Bo Pang, Zhenhai Zhu +1
Instructional videos get high-traffic on video sharing platforms, and prior work suggests that providing time-stamped, subtask annotations (e.g., "heat the oil in the pan") improve…
Understanding Image and Text Simultaneously: a Dual Vision-Language Machine Comprehension Task
Nan Ding, Sebastian Goodman, Fei Sha +1
We introduce a new multi-modal task for computer systems, posed as a combined vision-language comprehension challenge: identifying the most suitable text describing a scene, given…
Multilingual Word Embeddings using Multigraphs
Radu Soricut, Nan Ding
We present a family of neural-network--inspired models for computing continuous word representations, specifically designed to exploit both monolingual and multilingual text. This…
Building Large Machine Reading-Comprehension Datasets using Paragraph Vectors
Radu Soricut, Nan Ding
We present a dual contribution to the task of machine reading-comprehension: a technique for creating large-sized machine-comprehension (MC) datasets using paragraph-vector models;…