1 paper
Baptiste Rossigneux, Inna Kucher, Vincent Lorrain +1
Recent Visual-Language Models (VLMs) have enhanced the capabilities of pre-trained LLMs by adding vision tokens alongside text, with approaches like LLaVA showing impressive result…