5 papers · 1 filter
Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
Ting-Hao 'Kenneth' Huang, Ryan A. Rossi, Sungchul Kim +5
Between 2021 and 2025, the SciCap project grew from a small seed-funded idea at The Pennsylvania State University (Penn State) into one of the central efforts shaping the scientifi…
VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-use
Zhehao Zhang, Ryan Rossi, Tong Yu +7
While vision-language models (VLMs) have demonstrated remarkable performance across various tasks combining textual and visual information, they continue to struggle with fine-grai…
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
Ho Yin 'Sam' Ng, Ting-Yao Hsu, Aashish Anantha Ramakrishnan +8
Figure captions are crucial for helping readers understand and remember a figure's key message. Many models have been developed to generate these captions, helping authors compose…
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
Ting-Yao E. Hsu, Yi-Li Hsu, Shaurya Rohatgi +8
Since the SciCap datasets launch in 2021, the research community has made significant progress in generating captions for scientific figures in scholarly articles. In 2023, the fir…
Multi-LLM Collaborative Caption Generation in Scientific Documents
Jaeyoung Kim, Jongho Lee, Hong-Jun Choi +8
Scientific figure captioning is a complex task that requires generating contextually appropriate descriptions of visual content. However, existing methods often fall short by utili…