3 papers
cs.CV2026
PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language
Ijazul Haq, Yingjie Zhang, Irfan Ali Khan
This paper evaluates the performance of Large Multimodal Models (LMMs) on Optical Character Recognition (OCR) in the low-resource Pashto language. Natural Language Processing (NLP)…
cs.CV2025
A benchmark multimodal oro-dental dataset for large vision-language models
Haoxin Lv, Ijazul Haq, Jin Du +7
The advancement of artificial intelligence in oral healthcare relies on the availability of large-scale multimodal datasets that capture the complexity of clinical practice. In thi…
cs.CV2025
Transformer-based Spatial Grounding: A Comprehensive Survey
Ijazul Haq, Muhammad Saqib, Yingjie Zhang
Spatial grounding, the process of associating natural language expressions with corresponding image regions, has rapidly advanced due to the introduction of transformer-based model…