Evaluation of Automatic Video Captioning Using Direct Assessment
arXiv:1710.10586 · doi:10.1371/journal.pone.0202789
Abstract
We present Direct Assessment, a method for manually assessing the quality of automatically-generated captions for video. Evaluating the accuracy of video captions is particularly difficult because for any given video clip there is no definitive ground truth or correct answer against which to measure. Automatic metrics for comparing automatic video captions against a manual caption such as BLEU and METEOR, drawn from techniques used in evaluating machine translation, were used in the TRECVid video captioning task in 2016 but these are shown to have weaknesses. The work presented here brings human assessment into the evaluation by crowdsourcing how well a caption describes a video. We automatically degrade the quality of some sample captions which are assessed manually and from this we are able to rate the quality of the human assessors, a factor we take into account in the evaluation. Using data from the TRECVid video-to-text task in 2016, we show how our direct assessment method is replicable and robust and should scale to where there many caption-generation techniques to be evaluated.
26 pages, 8 figures
Cited by in corpus (4)
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- TRECVID 2020: A comprehensive campaign for evaluating video retrieval tasks across multiple application domains
- Video captioning with stacked attention and semantic hard pull
- On conducting better validation studies of automatic metrics in natural language generation evaluation