collaborators

5 papers

cs.CV2026

Do Composed Image Retrieval Benchmarks Require Multimodal Composition?

Matteo Attimonelli, Alessandro De Bellis, Aryo Pradipta Gema +8

Composed Image Retrieval (CIR) is a multimodal retrieval task where a query consists of a reference image and a textual modification, and the goal is to retrieve a target image sat…

cs.LG2026

Finding Culture-Sensitive Neurons in Vision-Language Models

Xiutian Zhao, Rochelle Choenni, Rohit Saxena +1

Despite their impressive performance, vision-language models (VLMs) still struggle on culturally situated inputs. To understand how VLMs process culturally grounded information, we…

cs.CV2026

VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models

Rohit Saxena, Alessandro Suglia, Pasquale Minervini

Vision-language models (VLMs) achieve strong performance on standard, high-quality datasets, but we still do not fully understand how they perform under real-world image distortion…

cs.CL2026

Analyzing LLM Instruction Optimization for Tabular Fact Verification

Xiaotang Du, Giwon Hong, Wai-Chung Kwan +4

Instruction optimization provides a lightweight, model-agnostic approach to enhancing the reasoning performance of large language models (LLMs). This paper presents the first syste…

cs.CL2025

Enhancing Long Document Long Form Summarisation with Self-Planning

Xiaotang Du, Rohit Saxena, Laura Perez-Beltrachini +2

We introduce a novel approach for long context summarisation, highlight-guided generation, that leverages sentence-level information as a content plan to improve the traceability a…