5 papers
Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?
Tawsif Tashwar Dipto, Azmol Hossain, Rubayet Sabbir Faruque +12
Conventional research on speech recognition modeling relies on the canonical form for most low-resource languages while automatic speech recognition (ASR) for regional dialects is…
Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
Shimanto Bhowmik, Tawsif Tashwar Dipto, Md Sazzad Islam +2
Bengali is an underrepresented language in NLP research. However, it remains a challenge due to its unique linguistic structure and computational constraints. In this work, we syst…
MSTT-199: MRI Dataset for Musculoskeletal Soft Tissue Tumor Segmentation
Tahsin Reasat, Stephen Chenard, Akhil Rekulapelli +5
Accurate musculoskeletal soft tissue tumor segmentation is vital for assessing tumor size, location, diagnosis, and response to treatment, thereby influencing patient outcomes. How…
Data Efficient Contrastive Learning in Histopathology using Active Sampling
Tahsin Reasat, Asif Sushmit, David S. Smith
Deep learning (DL) based diagnostics systems can provide accurate and robust quantitative analysis in digital pathology. These algorithms require large amounts of annotated trainin…
Unicode Normalization and Grapheme Parsing of Indic Languages
Nazmuddoha Ansary, Quazi Adibur Rahman Adib, Tahsin Reasat +6
Writing systems of Indic languages have orthographic syllables, also known as complex graphemes, as unique horizontal units. A prominent feature of these languages is these complex…