5 papers
Action-guided generation of 3D functionality segmentation data
Jaime Corsetti, Francesco Giuliari, Davide Boscaini +6
3D functionality segmentation aims to identify the interactive element in a 3D scene required to perform an action described in free-form language (e.g., the handle to ``Open the s…
Obstruction reasoning for robotic grasping
Runyu Jiao, Matteo Bortolon, Francesco Giuliari +5
Successful robotic grasping in cluttered environments not only requires a model to visually ground a target object but also to reason about obstructions that must be cleared before…
Free-form language-based robotic reasoning and grasping
Runyu Jiao, Alice Fasoli, Francesco Giuliari +5
Performing robotic grasping from a cluttered bin based on human instructions is a challenging task, as it requires understanding both the nuances of free-form language and the spat…
An analysis of vision-language models for fabric retrieval
Francesco Giuliari, Asif Khan Pattan, Mohamed Lamine Mekhalfi +1
Effective cross-modal retrieval is essential for applications like information retrieval and recommendation systems, particularly in specialized domains such as manufacturing, wher…
Functionality understanding and segmentation in 3D scenes
Jaime Corsetti, Francesco Giuliari, Alice Fasoli +2
Understanding functionalities in 3D scenes involves interpreting natural language descriptions to locate functional interactive objects, such as handles and buttons, in a 3D enviro…