1 citations · 1 across the 3 of their papers we have counts for
3 papers
Naïve PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation
Joong Ho Kim, Nicholas Thai, Souhardya Saha Dip +2
Text-to-Image (T2I) generation is primarily driven by Diffusion Models (DM) which rely on random Gaussian noise. Thus, like playing the slots at a casino, a DM will produce differe…
OOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking
Fazle Rabbi Rakib, Souhardya Saha Dip, Samiul Alam +11
We present OOD-Speech, the first out-of-distribution (OOD) benchmarking dataset for Bengali automatic speech recognition (ASR). Being one of the most spoken languages globally, Ben…
BaDLAD: A Large Multi-Domain Bengali Document Layout Analysis Dataset
Md. Istiak Hossain Shihab, Md. Rakibul Hasan, Mahfuzur Rahman Emon +14
While strides have been made in deep learning based Bengali Optical Character Recognition (OCR) in the past decade, the absence of large Document Layout Analysis (DLA) datasets has…