10 citations · 13 across the 8 of their papers we have counts for
4 papers · 1 filter
Unicode Normalization and Grapheme Parsing of Indic Languages
Nazmuddoha Ansary, Quazi Adibur Rahman Adib, Tahsin Reasat +6
Writing systems of Indic languages have orthographic syllables, also known as complex graphemes, as unique horizontal units. A prominent feature of these languages is these complex…
OOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking
Fazle Rabbi Rakib, Souhardya Saha Dip, Samiul Alam +11
We present OOD-Speech, the first out-of-distribution (OOD) benchmarking dataset for Bengali automatic speech recognition (ASR). Being one of the most spoken languages globally, Ben…
Data Efficient Contrastive Learning in Histopathology using Active Sampling
Tahsin Reasat, Asif Sushmit, David S. Smith
Deep learning (DL) based diagnostics systems can provide accurate and robust quantitative analysis in digital pathology. These algorithms require large amounts of annotated trainin…
BaDLAD: A Large Multi-Domain Bengali Document Layout Analysis Dataset
Md. Istiak Hossain Shihab, Md. Rakibul Hasan, Mahfuzur Rahman Emon +14
While strides have been made in deep learning based Bengali Optical Character Recognition (OCR) in the past decade, the absence of large Document Layout Analysis (DLA) datasets has…