4 papers
Exploring Cross-Lingual Knowledge Transfer via Transliteration-Based MLM Fine-Tuning for Critically Low-resource Chakma Language
Adity Khisa, Nusrat Jahan Lia, Tasnim Mahfuz Nafis +4
As an Indo-Aryan language with limited available data, Chakma remains largely underrepresented in language models. In this work, we introduce a novel corpus of contextually coheren…
Stemming -- The Evolution and Current State with a Focus on Bangla
Abhijit Paul, Mashiat Amin Farin, Sharif Md. Abdullah +3
Bangla, the seventh most widely spoken language worldwide with 300 million native speakers, faces digital under-representation due to limited resources and lack of annotated datase…
Where Journalism Silenced Voices: Exploring Discrimination in the Representation of Indigenous Communities in Bangladesh
Abhijit Paul, Adity Khisa, Zarif Masud +3
In this paper, we examine the intersections of indigeneity and media representation in shaping perceptions of indigenous communities in Bangladesh. Using a mixed-methods approach,…
The Rise of Small Language Models in Healthcare: A Comprehensive Survey
Muskan Garg, Shaina Raza, Shebuti Rayana +2
Despite substantial progress in healthcare applications driven by large language models (LLMs), growing concerns around data privacy, and limited resources; the small language mode…