6 citations · 24 across the 39 of their papers we have counts for
5 papers · 1 filter
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
Rohit Saxena, Alessandro Suglia, Pasquale Minervini
Vision-language models (VLMs) achieve strong performance on standard, high-quality datasets, but we still do not fully understand how they perform under real-world image distortion…
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
Rohit Saxena, Aryo Pradipta Gema, Pasquale Minervini
Understanding time from visual representations is a fundamental cognitive skill, yet it remains a challenge for multimodal large language models (MLLMs). In this work, we investiga…
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
Rohit Saxena, Pasquale Minervini, Frank Keller
Generating accurate and concise textual summaries from multimodal documents is challenging, especially when dealing with visually complex content like scientific posters. We introd…
With Great Backbones Comes Great Adversarial Transferability
Erik Arakelyan, Karen Hambardzumyan, Davit Papikyan +4
Advances in self-supervised learning (SSL) for machine vision have improved representation robustness and model performance, giving rise to pre-trained backbones like \emph{ResNet}…
Adaptive Layer Selection for Efficient Vision Transformer Fine-Tuning
Alessio Devoto, Federico Alvetreti, Jary Pomponi +3
Recently, foundation models based on Vision Transformers (ViTs) have become widely available. However, their fine-tuning process is highly resource-intensive, and it hinders their…