1 paper
Ana Carolina Condez, Diogo Tavares, João Magalhães
Recent advances in vision-language models have enabled rich semantic understanding across modalities. However, these encoding methods lack the ability to interpret or reason about…