1 paper
Maksim Dzabraev, Alexander Kunitsyn, Andrei Ivaniuta
In this work, we present an unsupervised method for enhancing an image captioning model (in our case, BLIP2) using reinforcement learning and vision-language models like CLIP and B…