2 papers
cs.CV2026
An overview of 3D Vision-Language Models
Márcus Lobo, Vitor Matias, Afonso Paiva +3
Vision-Language Models (VLMs) are reshaping computer vision by aligning visual and textual embeddings, allowing models to recognize visual concepts and reason about them using natu…
cs.CV2026
3D-MRL: Nested Multimodal 3D Representations via Matryoshka Representation Learning
Márcus Lobo, Vitor Matias, Jeová Farias +1
Vision-Language Models align point clouds with image and text embeddings, enabling zero-shot recognition, retrieval, and open-vocabulary understanding of 3D shapes. Existing multim…