1 paper
Jakob Krogh Petersen, Valdemar Licht, Mads Nielsen +1
Multi-modal models require aligned, shared embedding spaces. However, common CLIP-based approaches need large amounts of samples and do not natively support 3D or tabular data, bot…