9 papers
Action-guided generation of 3D functionality segmentation data
Jaime Corsetti, Francesco Giuliari, Davide Boscaini +6
3D functionality segmentation aims to identify the interactive element in a 3D scene required to perform an action described in free-form language (e.g., the handle to ``Open the s…
Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
Guofeng Mei, Wei Lin, Luigi Riz +3
Large Multimodal Models (LMMs) that process 3D data typically rely on heavy, pre-trained visual encoders to extract geometric features. While recent 2D LMMs have begun to eliminate…
Obstruction reasoning for robotic grasping
Runyu Jiao, Matteo Bortolon, Francesco Giuliari +5
Successful robotic grasping in cluttered environments not only requires a model to visually ground a target object but also to reason about obstructions that must be cleared before…
OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields
Lisa Weijler, Sebastian Koch, Fabio Poiesi +2
Modeling the inherent hierarchical structure of 3D objects and 3D scenes is highly desirable, as it enables a more holistic understanding of environments for autonomous agents. Acc…
Free-form language-based robotic reasoning and grasping
Runyu Jiao, Alice Fasoli, Francesco Giuliari +5
Performing robotic grasping from a cluttered bin based on human instructions is a challenging task, as it requires understanding both the nuances of free-form language and the spat…
PerLA: Perceptive 3D Language Assistant
Guofeng Mei, Wei Lin, Luigi Riz +3
Enabling Large Language Models (LLMs) to understand the 3D physical world is an emerging yet challenging research direction. Current strategies for processing point clouds typicall…