3 papers
cs.CV2020
Enriching Video Captions With Contextual Text
Philipp Rimle, Pelin Dogan, Markus Gross
Understanding video content and generating caption with context is an important and challenging task. Unlike prior methods that typically attempt to generate generic video captions…
cs.CV2019
Neural Sequential Phrase Grounding (SeqGROUND)
Pelin Dogan, Leonid Sigal, Markus Gross
We propose an end-to-end approach for phrase grounding in images. Unlike prior methods that typically attempt to ground each phrase independently by building an image-text embeddin…
cs.CV2018
A Neural Multi-sequence Alignment TeCHnique (NeuMATCH)
Pelin Dogan, Boyang Li, Leonid Sigal +1
The alignment of heterogeneous sequential data (video to text) is an important and challenging problem. Standard techniques for this task, including Dynamic Time Warping (DTW) and…