10 papers
Federated Large Language Models: Current Progress and Future Directions
Yuhang Yao, Jianyi Zhang, Junda Wu +11
Large Language Models have achieved impressive performance across diverse applications, yet their training typically depends on centralized data collection, raising serious privacy…
Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
Ting-Hao 'Kenneth' Huang, Ryan A. Rossi, Sungchul Kim +5
Between 2021 and 2025, the SciCap project grew from a small seed-funded idea at The Pennsylvania State University (Penn State) into one of the central efforts shaping the scientifi…
Charts Are Not Images: On the Challenges of Scientific Chart Editing
Shawn Li, Ryan Rossi, Sungchul Kim +5
Generative models, such as diffusion and autoregressive approaches, have demonstrated impressive capabilities in editing natural images. However, applying these tools to scientific…
VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-use
Zhehao Zhang, Ryan Rossi, Tong Yu +7
While vision-language models (VLMs) have demonstrated remarkable performance across various tasks combining textual and visual information, they continue to struggle with fine-grai…
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
Ho Yin 'Sam' Ng, Ting-Yao Hsu, Aashish Anantha Ramakrishnan +8
Figure captions are crucial for helping readers understand and remember a figure's key message. Many models have been developed to generate these captions, helping authors compose…
WaterFlow: Learning Fast & Robust Watermarks using Stable Diffusion
Vinay Shukla, Prachee Sharma, Ryan Rossi +3
The ability to embed watermarks in images is a fundamental problem of interest for computer vision, and is exacerbated by the rapid rise of generated imagery in recent times. Curre…