4 papers
Towards Detecting Inconsistencies in End-to-end Generated TODs
Tiziano Labruna, Giovanni Bonetta, Bernardo Magnini
Generative AI is profoundly transforming the core technologies behind conversational systems, shifting from component-based to end-to-end approaches. However, Large Language Models…
ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models
Matteo Merler, Nicola Dainese, Minttu Alakuijala +5
Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans, with recent work extending this idea to visual domain…
All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark
Davide Testa, Giovanni Bonetta, Raffaella Bernardi +5
We introduce MAIA (Multimodal AI Assessment), a native-Italian benchmark designed for fine-grained investigation of the reasoning abilities of visual language models on videos. MAI…
Evaluating Task-Oriented Dialogue Consistency through Constraint Satisfaction
Tiziano Labruna, Bernardo Magnini
Task-oriented dialogues must maintain consistency both within the dialogue itself, ensuring logical coherence across turns, and with the conversational domain, accurately reflectin…