1 paper · 1 filter
Yiwei Chen, Yuguang Yao, Yihua Zhang +3
Recent vision language models (VLMs) have made remarkable strides in generative modeling with multimodal inputs, particularly text and images. However, their susceptibility to gene…