1 paper
Daniel Shalam, Emanuel Ben Baruch, Avi Ben Cohen +1
Multimodal large language models can emit localized predictions, bounding boxes for objects and temporal windows for video and audio events, but they hallucinate these regions prol…