3 papers
cs.CV2025
EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models
Haiyang Yu, Mengyang Zhao, Jinghui Lu +8
Video subtitles play a crucial role in short videos and movies, as they not only help models better understand video content but also support applications such as video translation…
cs.CL2025
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
Haiyang Yu, Yuchuan Wu, Fan Shi +18
Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and unders…
cs.CV2025
UMIT: Unifying Medical Imaging Tasks via Vision-Language Models
Haiyang Yu, Siyang Yi, Ke Niu +2
With the rapid advancement of deep learning, particularly in the field of medical image analysis, an increasing number of Vision-Language Models (VLMs) are being widely applied to…