1 paper · 1 filter
Jawad Ibn Ahad, Maisha Rahman, Amrijit Biswas +5
Multimodal language models (MLLMs) require large parameter capacity to align high-dimensional visual features with linguistic representations, making them computationally heavy and…