5 papers
Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
Guofeng Mei, Wei Lin, Luigi Riz +3
Large Multimodal Models (LMMs) that process 3D data typically rely on heavy, pre-trained visual encoders to extract geometric features. While recent 2D LMMs have begun to eliminate…
Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization
Guofeng Mei, Bin Ren, Qinfeng Xiao +8
Vision-language models, such as CLIP, encode rich semantic knowledge through large-scale image-text pretraining. Reusing these models for 3D understanding is highly desirable, beca…
PerLA: Perceptive 3D Language Assistant
Guofeng Mei, Wei Lin, Luigi Riz +3
Enabling Large Language Models (LLMs) to understand the 3D physical world is an emerging yet challenging research direction. Current strategies for processing point clouds typicall…
Vocabulary-Free 3D Instance Segmentation with Vision and Language Assistant
Guofeng Mei, Luigi Riz, Yiming Wang +1
Most recent 3D instance segmentation methods are open vocabulary, offering a greater flexibility than closed-vocabulary methods. Yet, they are limited to reasoning within a specifi…
Wild Berry image dataset collected in Finnish forests and peatlands using drones
Luigi Riz, Sergio Povoli, Andrea Caraffa +13
Berry picking has long-standing traditions in Finland, yet it is challenging and can potentially be dangerous. The integration of drones equipped with advanced imaging techniques r…