4 papers
Humans vs Vision-Language Models: A Unified Measure of Narrative Coherence
Nikolai Ilinykh, Hyewon Jang, Shalom Lappin +2
We study narrative coherence in visually grounded stories by comparing human-written narratives with those generated by vision-language models (VLMs) on the Visual Writing Prompts…
Predicting Sentence Acceptability Judgments in Multimodal Contexts
Hyewon Jang, Nikolai Ilinykh, Sharid Loáiciga +2
Previous work has examined the capacity of deep neural networks (DNNs), particularly transformers, to predict human sentence acceptability judgments, both independently of context,…
Surprisal reveals diversity gaps in image captioning and different scorers change the story
Nikolai Ilinykh, Simon Dobnik
We quantify linguistic diversity in image captioning with surprisal variance - the spread of token-level negative log-probabilities within a caption set. On the MSCOCO test set, we…
Coreference as an indicator of context scope in multimodal narrative
Nikolai Ilinykh, Shalom Lappin, Asad Sayeed +1
We demonstrate that large multimodal language models differ substantially from humans in the distribution of coreferential expressions in a visual storytelling task. We introduce a…