1 paper
Yuhan Chen, Lumei Su, Lihua Chen +1
In this paper, the LCV2 modular method is proposed for the Grounded Visual Question Answering task in the vision-language multimodal domain. This approach relies on a frozen large…