1 paper · 1 filter
Armin Gerami, Seyedehanita Madani, Ramani Duraiswami
Multimodal Transformers serve as the backbone for state-of-the-art vision-language models, yet their quadratic attention complexity remains a critical barrier to scalability. In th…