Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models
Haiyang Yu, Mengyang Zhao, Jinghui Lu +8
Video subtitles play a crucial role in short videos and movies, as they not only help models better understand video content but also support applications such as video translation…
cs.CV2025
UMIT: Unifying Medical Imaging Tasks via Vision-Language Models
Haiyang Yu, Siyang Yi, Ke Niu +2
With the rapid advancement of deep learning, particularly in the field of medical image analysis, an increasing number of Vision-Language Models (VLMs) are being widely applied to…