1 paper
Riad Ahmed Anonto, Sardar Md. Saffat Zabin, M. Saifur Rahman
Grounding vision--language models in low-resource languages remains challenging, as they often produce fluent text about the wrong objects. This stems from scarce paired data, tran…