11 papers
GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes
Amit Parekh, Sabrina McCallum, Kareem Al-Hasan +3
Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents. Existing benchmarks show that these models possess many of the re…
Retrievit: In-context Retrieval Capabilities of Transformers, State Space Models, and Hybrid Architectures
Georgios Pantazopoulos, Malvina Nikandrou, Ioannis Konstas +1
Transformers excel at in-context retrieval but suffer from quadratic complexity with sequence length, while State Space Models (SSMs) offer efficient linear-time processing but hav…
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
Matteo Attimonelli, Alessandro De Bellis, Aryo Pradipta Gema +8
Composed Image Retrieval (CIR) is a multimodal retrieval task where a query consists of a reference image and a textual modification, and the goal is to retrieve a target image sat…
MIXAR: Scaling Autoregressive Pixel-based Language Models to Multiple Languages and Scripts
Chen Hu, Yintao Tai, Antonio Vergari +2
Pixel-based language models are gaining momentum as alternatives to traditional token-based approaches, promising to circumvent tokenization challenges. However, the inherent perce…
AgriPath: A Systematic Exploration of Architectural Trade-offs for Crop Disease Classification
Hamza Mooraj, George Pantazopoulos, Alessandro Suglia
Reliable crop disease detection requires models that perform consistently across diverse acquisition conditions, yet existing evaluations often focus on single architectural famili…
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
Rohit Saxena, Alessandro Suglia, Pasquale Minervini
Vision-language models (VLMs) achieve strong performance on standard, high-quality datasets, but we still do not fully understand how they perform under real-world image distortion…