5 papers
Reproducible Multimodal Affordance Prediction
Tommaso Apicella, Alessio Xompero, Andrea Cavallaro
Affordance prediction is the identification of potential actions an agent can perform on a target object from multimodal inputs. Affordance prediction methods are difficult to eval…
Visual Affordance Prediction: Survey and Reproducibility
Tommaso Apicella, Alessio Xompero, Andrea Cavallaro
Affordances are the potential actions an agent can perform on an object, as observed by a camera. Visual affordance prediction is formulated differently for tasks such as grasping…
On the Robustness of Vision-Language Models in Zero-shot Privacy Classification
Alina Elena Baia, Alessio Xompero, Andrea Cavallaro
Automatic systems for document understanding require multimodal models that accurately identify sensitive visual content, even in the presence of image degradations. Instruction-fo…
Stereo Hand-Object Reconstruction for Human-to-Robot Handover
Yik Lung Pang, Alessio Xompero, Changjae Oh +1
Jointly estimating hand and object shape facilitates the grasping task in human-to-robot handovers. However, relying on hand-crafted prior knowledge about the geometric structure o…
Learning Privacy from Visual Entities
Alessio Xompero, Andrea Cavallaro
Subjective interpretation and content diversity make predicting whether an image is private or public a challenging task. Graph neural networks combined with convolutional neural n…