1 citations · 1 across the 1 of their papers we have counts for
1 paper
Shayon Dasgupta, Avijit Dasgupta, C. V. Jawahar
Multimodal Large Language Models (MLLMs) have redefined visual understanding by combining vision encoders with large-scale language models. This unified architecture enables strong…