2 papers
cs.CV2026
CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
Kumal Hewagamage, Isuranga Senavirathne, Sasika Amarasinghe +4
4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vision encoders are largely limited…
cs.CV2025
CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images
Avishka Perera, Kumal Hewagamage, Saeedha Nazar +4
Image-to-point cross-modal learning has emerged to address the scarcity of large-scale 3D datasets in 3D representation learning. However, current methods that leverage 2D data oft…