2 papers
cs.CV2025
AutoArabic: A Three-Stage Framework for Localizing Video-Text Retrieval Benchmarks
Mohamed Eltahir, Osamah Sarraj, Abdulrahman Alfrihidi +4
Video-to-text and text-to-video retrieval are dominated by English benchmarks (e.g. DiDeMo, MSR-VTT) and recent multilingual corpora (e.g. RUDDER), yet Arabic remains underserved,…
cs.CV2025
Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric
Mohamed Eltahir, Osamah Sarraj, Mohammed Bremoo +5
Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively wit…