1 citations · 1 across the 2 of their papers we have counts for
3 papers · 1 filter
BanglaParaphrase: A High-Quality Bangla Paraphrase Dataset
Ajwad Akil, Najrin Sultana, Abhik Bhattacharjee +1
In this work, we present BanglaParaphrase, a high-quality synthetic Bangla Paraphrase dataset curated by a novel filtering pipeline. We aim to take a step towards alleviating the l…
XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages
Tahmid Hasan, Abhik Bhattacharjee, Md Saiful Islam +5
Contemporary works on abstractive text summarization have focused primarily on high-resource languages like English, mostly due to the limited availability of datasets for low/mid-…
Not Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for Bengali-English Machine Translation
Tahmid Hasan, Abhik Bhattacharjee, Kazi Samin +4
Despite being the seventh most widely spoken language in the world, Bengali has received much less attention in machine translation literature due to being low in resources. Most p…