activity
20242026
collaborators

6 papers

cs.CL2026

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

Davide Testa, Hugh Mee Wong, Alessandro Lenci +2

Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual…

cs.AI2026

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

Giovanni Bonetta, Matteo Merler, Davide Zago +2

Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as policies is expensive and brittle: they must be queried at every ste…

cs.CL2026

Towards Detecting Inconsistencies in End-to-end Generated TODs

Tiziano Labruna, Giovanni Bonetta, Bernardo Magnini

Generative AI is profoundly transforming the core technologies behind conversational systems, shifting from component-based to end-to-end approaches. However, Large Language Models…

cs.AI2025

ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models

Matteo Merler, Nicola Dainese, Minttu Alakuijala +5

Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans, with recent works extending this idea to visual domai…

cs.CL2025

All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark

Davide Testa, Giovanni Bonetta, Raffaella Bernardi +5

We introduce MAIA (Multimodal AI Assessment), a native-Italian benchmark designed for fine-grained investigation of the reasoning abilities of visual language models on videos. MAI…

cs.CL2024

Evaluating Task-Oriented Dialogue Consistency through Constraint Satisfaction

Tiziano Labruna, Bernardo Magnini

Task-oriented dialogues must maintain consistency both within the dialogue itself, ensuring logical coherence across turns, and with the conversational domain, accurately reflectin…