5 papers
Breaking the Silence: A Dataset and Benchmark for Bangla Text-to-Gloss Translation
Sharif Mohammad Abdullah, Abhijit Paul, Shubhashis Roy Dipta +3
Gloss is a written approximation that bridges Sign Language (SL) and its corresponding spoken language. Despite a deaf and hard-of-hearing population of at least 3 million in Bangl…
Exploring Cross-Lingual Knowledge Transfer via Transliteration-Based MLM Fine-Tuning for Critically Low-resource Chakma Language
Adity Khisa, Nusrat Jahan Lia, Tasnim Mahfuz Nafis +4
As an Indo-Aryan language with limited available data, Chakma remains largely underrepresented in language models. In this work, we introduce a novel corpus of contextually coheren…
Stemming -- The Evolution and Current State with a Focus on Bangla
Abhijit Paul, Mashiat Amin Farin, Sharif Md. Abdullah +3
Bangla, the seventh most widely spoken language worldwide with 300 million native speakers, faces digital under-representation due to limited resources and lack of annotated datase…
Where Journalism Silenced Voices: Exploring Discrimination in the Representation of Indigenous Communities in Bangladesh
Abhijit Paul, Adity Khisa, Zarif Masud +3
In this paper, we examine the intersections of indigeneity and media representation in shaping perceptions of indigenous communities in Bangladesh. Using a mixed-methods approach,…
The Rise of Small Language Models in Healthcare: A Comprehensive Survey
Muskan Garg, Shaina Raza, Shebuti Rayana +2
Despite substantial progress in healthcare applications driven by large language models (LLMs), growing concerns around data privacy, and limited resources; the small language mode…