2 papers
cs.CV2025
EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models
Haiyang Yu, Mengyang Zhao, Jinghui Lu +8
Video subtitles play a crucial role in short videos and movies, as they not only help models better understand video content but also support applications such as video translation…
cs.CV2025
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
Yuxiang Nie, Han Wang, Yongjie Ye +15
This paper introduces ChineseVideoBench, a pioneering benchmark specifically designed for evaluating Multimodal Large Language Models (MLLMs) in Chinese Video Question Answering. T…