2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CL2025
GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art
Yiming Lei, Chenkai Zhang, Zeming Liu +5
Video Comment Art enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural…
cs.CV2025
SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
Chenkai Zhang, Yiming Lei, Zeming Liu +5
With the rapid development of Multi-modal Large Language Models (MLLMs), an increasing number of benchmarks have been established to evaluate the video understanding capabilities o…
cs.CV2025★ 2 cited
A Survey on Remote Sensing Foundation Models: From Vision to Multimodality
Ziyue Huang, Hongxi Yan, Qiqi Zhan +7
The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of intelligent geospatial data int…