8 papers
Assessing the Visual Enumeration Abilities of Specialized Counting Architectures and Vision-Language Models
Kuinan Hou, Jing Mi, Marco Zorzi +2
Counting the number of items in a visual scene remains a fundamental yet challenging task in computer vision. Traditional approaches to solving this problem rely on domain-specific…
PersONAL: Towards a Comprehensive Benchmark for Personalized Embodied Agents
Filippo Ziliotto, Jelin Raphael Akkara, Alessandro Daniele +3
Recent advances in Embodied AI have enabled agents to perform increasingly complex tasks and adapt to diverse environments. However, deploying such agents in realistic human-center…
7Bench: a Comprehensive Benchmark for Layout-guided Text-to-image Models
Elena Izzo, Luca Parolari, Davide Vezzaro +1
Layout-guided text-to-image models offer greater control over the generation process by explicitly conditioning image synthesis on the spatial arrangement of elements. As a result,…
Temporally-Aware Supervised Contrastive Learning for Polyp Counting in Colonoscopy
Luca Parolari, Andrea Cherubini, Lamberto Ballan +1
Automated polyp counting in colonoscopy is a crucial step toward automated procedure reporting and quality control, aiming to enhance the cost-effectiveness of colonoscopy screenin…
MLFM: Multi-Layered Feature Maps for Richer Language Understanding in Zero-Shot Semantic Navigation
Sonia Raychaudhuri, Enrico Cancelli, Tommaso Campari +3
Recent progress in large vision-language models has driven improvements in language-based semantic navigation, where an embodied agent must reach a target object described in natur…
Towards Polyp Counting In Full-Procedure Colonoscopy Videos
Luca Parolari, Andrea Cherubini, Lamberto Ballan +1
Automated colonoscopy reporting holds great potential for enhancing quality control and improving cost-effectiveness of colonoscopy procedures. A major challenge lies in the automa…