1 paper
Luigi Celona, Simone Bianco, Marco Donzella +1
State-of-The-Art (SoTA) image captioning models are often trained on the MicroSoft Common Objects in Context (MS-COCO) dataset, which contains human-annotated captions with an aver…