1 paper · 1 filter
Deqing Fu, Ruohao Guo, Ghazal Khalighinejad +5
Current foundation models exhibit impressive capabilities when prompted either with text only or with both image and text inputs. But do their capabilities change depending on the…