3 papers
cs.CL2025
Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration
Kanchon Gharami, Quazi Sarwar Muhtaseem, Deepti Gupta +2
The development of robust transliteration techniques to enhance the effectiveness of transforming Romanized scripts into native scripts is crucial for Natural Language Processing t…
cs.CL2025
TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking
Shahriar Kabir Nahin, Rabindra Nath Nandi, Sagor Sarker +7
In this paper, we present TituLLMs, the first large pretrained Bangla LLMs, available in 1b and 3b parameter sizes. Due to computational constraints during both training and infere…
cs.CL2023
Pseudo-Labeling for Domain-Agnostic Bangla Automatic Speech Recognition
Rabindra Nath Nandi, Mehadi Hasan Menon, Tareq Al Muntasir +5
One of the major challenges for developing automatic speech recognition (ASR) for low-resource languages is the limited access to labeled data with domain-specific variations. In t…