3 papers
cs.CL2025
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
Haneen Al-Homoud, Asma Ibrahim, Murtadha Al-Jubran +5
Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (D…
cs.CV2025
Saudi Sign Language Translation Using T5
Ali Alhejab, Tomas Zelezny, Lamya Alkanhal +7
This paper explores the application of T5 models for Saudi Sign Language (SSL) translation using a novel dataset. The SSL dataset includes three challenging testing protocols, enab…
cs.CV2025
Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition
Sarah Alyami, Hamzah Luqman, Sadam Al-Azani +3
Current benchmarks for sign language recognition (SLR) focus mainly on isolated SLR, while there are limited datasets for continuous SLR (CSLR), which recognizes sequences of signs…