1 paper · 1 filter
Stanley Cao, Sonny Young
Image captioning using Vision Transformers (ViTs) represents a pivotal convergence of computer vision and natural language processing, offering the potential to enhance user experi…