most citedVisual Prompting in Multimodal Large Language Models: A Survey

4 citations · 4 across the 5 of their papers we have counts for

collaborators

5 papers

cs.IR2025

From Documents to Dialogue: Building KG-RAG Enhanced AI Assistants

Manisha Mukherjee, Sungchul Kim, Xiang Chen +3

The Adobe Experience Platform AI Assistant is a conversational tool that enables organizations to interact seamlessly with proprietary enterprise data through a chatbot. However, d…

cs.CL2025

Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023

Ting-Yao E. Hsu, Yi-Li Hsu, Shaurya Rohatgi +8

Since the SciCap datasets launch in 2021, the research community has made significant progress in generating captions for scientific figures in scholarly articles. In 2023, the fir…

cs.HC2025

Understanding How Paper Writers Use AI-Generated Captions in Figure Caption Writing

Ho Yin, Ng, Ting-Yao Hsu +6

Figures and their captions play a key role in scientific publications. However, despite their importance, many captions in published papers are poorly crafted, largely due to a lac…

cs.CL2025

Multi-LLM Collaborative Caption Generation in Scientific Documents

Jaeyoung Kim, Jongho Lee, Hong-Jun Choi +8

Scientific figure captioning is a complex task that requires generating contextually appropriate descriptions of visual content. However, existing methods often fall short by utili…

cs.LG20244 cited

Visual Prompting in Multimodal Large Language Models: A Survey

Junda Wu, Zhehao Zhang, Yu Xia +12

Multimodal large language models (MLLMs) equip pre-trained large-language models (LLMs) with visual capabilities. While textual prompting in LLMs has been widely studied, visual pr…