1 paper · 1 filter
Passant Elchafei, Amany Fashwan
We present VLCAP, an Arabic image captioning framework that integrates CLIP-based visual label retrieval with multimodal text generation. Rather than relying solely on end-to-end c…