1 paper
Prabal Shrestha, Bohan Jiang, Haoning Xue +2
Multimodal large language models (MLLMs) have shown strong performance on objective tasks such as video understanding and reasoning. However, it remains unclear whether they can ap…