collaborators

8 papers

cs.CV2025

Assessing the Visual Enumeration Abilities of Specialized Counting Architectures and Vision-Language Models

Kuinan Hou, Jing Mi, Marco Zorzi +2

Counting the number of items in a visual scene remains a fundamental yet challenging task in computer vision. Traditional approaches to solving this problem rely on domain-specific…

cs.CV2025

PersONAL: Towards a Comprehensive Benchmark for Personalized Embodied Agents

Filippo Ziliotto, Jelin Raphael Akkara, Alessandro Daniele +3

Recent advances in Embodied AI have enabled agents to perform increasingly complex tasks and adapt to diverse environments. However, deploying such agents in realistic human-center…

cs.CV2025

7Bench: a Comprehensive Benchmark for Layout-guided Text-to-image Models

Elena Izzo, Luca Parolari, Davide Vezzaro +1

Layout-guided text-to-image models offer greater control over the generation process by explicitly conditioning image synthesis on the spatial arrangement of elements. As a result,…

cs.CV2025

Temporally-Aware Supervised Contrastive Learning for Polyp Counting in Colonoscopy

Luca Parolari, Andrea Cherubini, Lamberto Ballan +1

Automated polyp counting in colonoscopy is a crucial step toward automated procedure reporting and quality control, aiming to enhance the cost-effectiveness of colonoscopy screenin…

cs.RO2025

MLFM: Multi-Layered Feature Maps for Richer Language Understanding in Zero-Shot Semantic Navigation

Sonia Raychaudhuri, Enrico Cancelli, Tommaso Campari +3

Recent progress in large vision-language models has driven improvements in language-based semantic navigation, where an embodied agent must reach a target object described in natur…

cs.CV2025

Towards Polyp Counting In Full-Procedure Colonoscopy Videos

Luca Parolari, Andrea Cherubini, Lamberto Ballan +1

Automated colonoscopy reporting holds great potential for enhancing quality control and improving cost-effectiveness of colonoscopy procedures. A major challenge lies in the automa…