1 paper
Waikit Xiu, Qiang Lu, Zian Wang +4
In safety-critical traffic scenarios, answering complex questions relies on minute, localized visual cues. However, standard Multimodal Large Language Models (MLLMs) tend to over-a…