19 citations · 46 across the 30 of their papers we have counts for
1 paper · 2 filters
Wish Suharitdamrong, Muhammad Awais, Xiatian Zhu +1
Multimodal Large Language Models (MLLMs) can listen and see, but how do audio and visual signals actually travel through the network to shape an answer? Despite their growing role…