1 paper
Ruoyu Zhang, Lulu Wang, Yi He +3
Recent advancements in large language models (LLMs) have significantly enhanced the fluency and logical coherence of image captioning. Retrieval-Augmented Generation (RAG) is widel…