1 paper · 1 filter
Zefan Cai, Haoyi Qiu, Tianyi Ma +13
Modern multimodal generative models can synthesize visually compelling images and videos, but it remains unclear whether this visual fluency reflects genuine reasoning: when prompt…