1 paper
Hyeonsu Kang, Emily Bao, Anjan Goswami
Vision-language models (VLMs) are increasingly used to evaluate multimodal content, including presentation slides, yet their slide-specific understanding remains underexplored {des…