2 papers
cs.CV2025
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
Xuyi Yang, Wenhao Zhang, Hongbo Jin +5
Current Multimodal Large Language Models (MLLMs) often perform poorly in long video understanding, primarily due to resource limitations that prevent them from processing all video…
cs.CL2025
EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models
He Hu, Lianzhong You, Hongbo Xu +7
With the integration of multimodal large language models (MLLMs) into robotic systems and AI applications, embedding emotional intelligence (EI) capabilities is essential for enabl…