1 paper · 1 filter
Bishwash Khanal, Anlan Zhang, Sasu Tarkoma +2
Vision-language models can describe images fluently, but they often fail to provide actionable photographic critique because semantic content and aesthetic judgment remain entangle…