Pre-gen metrics: Predicting caption quality metrics without generating captions
arXiv:1810.05474 · doi:10.1007/978-3-030-11018-5_11
Abstract
Image caption generation systems are typically evaluated against reference outputs. We show that it is possible to predict output quality without generating the captions, based on the probability assigned by the neural model to the reference captions. Such pre-gen metrics are strongly correlated to standard evaluation metrics.
13 pages, 6 figures This publication will appear in the Proceedings of the First Workshop on Shortcomings in Vision and Language (2018). DOI to be inserted later