2 papers
cs.CV2024
B-Cos Aligned Transformers Learn Human-Interpretable Features
Manuel Tran, Amal Lahiani, Yashin Dicente Cid +7
Vision Transformers (ViTs) and Swin Transformers (Swin) are currently state-of-the-art in computational pathology. However, domain experts are still reluctant to use these models d…
cs.AI2023
Training Transitive and Commutative Multimodal Transformers with LoReTTa
Manuel Tran, Yashin Dicente Cid, Amal Lahiani +3
Training multimodal foundation models is challenging due to the limited availability of multimodal datasets. While many public datasets pair images with text, few combine images wi…