7 papers · 1 filter
Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
Guofeng Mei, Wei Lin, Luigi Riz +3
Large Multimodal Models (LMMs) that process 3D data typically rely on heavy, pre-trained visual encoders to extract geometric features. While recent 2D LMMs have begun to eliminate…
Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization
Guofeng Mei, Qinfeng Xiao, Bin Ren +7
Vision-language models, such as CLIP, encode rich semantic knowledge through large-scale image-text pretraining. Reusing these models for 3D understanding is highly desirable, beca…
PerLA: Perceptive 3D Language Assistant
Guofeng Mei, Wei Lin, Luigi Riz +3
Enabling Large Language Models (LLMs) to understand the 3D physical world is an emerging yet challenging research direction. Current strategies for processing point clouds typicall…
Vocabulary-Free 3D Instance Segmentation with Vision and Language Assistant
Guofeng Mei, Luigi Riz, Yiming Wang +1
Most recent 3D instance segmentation methods are open vocabulary, offering a greater flexibility than closed-vocabulary methods. Yet, they are limited to reasoning within a specifi…
Wild Berry image dataset collected in Finnish forests and peatlands using drones
Luigi Riz, Sergio Povoli, Andrea Caraffa +13
Berry picking has long-standing traditions in Finland, yet it is challenging and can potentially be dangerous. The integration of drones equipped with advanced imaging techniques r…
Novel class discovery meets foundation models for 3D semantic segmentation
Luigi Riz, Cristiano Saltori, Yiming Wang +2
The task of Novel Class Discovery (NCD) in semantic segmentation entails training a model able to accurately segment unlabelled (novel) classes, relying on the available supervisio…