2 papers
cs.CV2025
A benchmark multimodal oro-dental dataset for large vision-language models
Haoxin Lv, Ijazul Haq, Jin Du +7
The advancement of artificial intelligence in oral healthcare relies on the availability of large-scale multimodal datasets that capture the complexity of clinical practice. In thi…
cs.CV2025
Transformer-based Spatial Grounding: A Comprehensive Survey
Ijazul Haq, Muhammad Saqib, Yingjie Zhang
Spatial grounding, the process of associating natural language expressions with corresponding image regions, has rapidly advanced due to the introduction of transformer-based model…