2 citations · 3 across the 2 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos
Zixi Yang, Jiapeng Li, Muxi Diao +2
Recently, Multi-modal Large Language Models (MLLMs) have demonstrated significant performance across various video understanding tasks. However, their robustness, particularly when…
cs.CV2024★ 1 cited
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
Hao Liang, Jiapeng Li, Tianyi Bai +7
Recently, with the rise of web videos, managing and understanding large-scale video datasets has become increasingly important. Video Large Language Models (VideoLLMs) have emerged…