1 paper
Ari Wahl, Dorian Gawlinski, David Przewozny +3
Pre-trained general-purpose Vision-Language Models (VLM) hold the potential to enhance intuitive human-machine interactions due to their rich world knowledge and 2D object detectio…