2 papers
cs.CV2025
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
Yuxuan Cai, Jiangning Zhang, Zhenye Gan +9
Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks involving both images and videos. However, their capacity to comprehen…
cs.CV2025
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
Yuxuan Cai, Jiangning Zhang, Haoyang He +7
The success of Large Language Models (LLMs) has inspired the development of Multimodal Large Language Models (MLLMs) for unified understanding of vision and language. However, the…