7 papers
Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering
Samir Wagle, Abiral Adhikari, Reewaj Khanal +4
Legal domains in high-resource languages like English have widely adopted artificial intelligence for legal question answering. However, data scarcity in low resource languages suc…
NwÄchÄ MunÄ: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
Rishikesh Kumar Sharma, Safal Narshing Shrestha, Jenny Poudel +4
Nepal Bhasha (Newari), an endangered language of the Kathmandu Valley, remains digitally marginalized due to the severe scarcity of annotated speech resources. In this work, we int…
NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
Rupak Raj Ghimire, Bipesh Subedi, Balaram Prasain +7
Modern Translation Systems heavily rely on high-quality, large parallel datasets for state-of-the-art performance. However, such resources are largely unavailable for most of the S…
Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications
Funghang Limbu Begha, Praveen Acharya, Bal Krishna Bal
Nepali, a low-resource language, faces significant challenges in building an effective information retrieval system due to the unavailability of annotated data and computational li…
Benchmarking BERT-based Models for Sentence-level Topic Classification in Nepali Language
Nischal Karki, Bipesh Subedi, Prakash Poudyal +2
Transformer-based models such as BERT have significantly advanced Natural Language Processing (NLP) across many languages. However, Nepali, a low-resource language written in Devan…
Consolidating and Developing Benchmarking Datasets for the Nepali Natural Language Understanding Tasks
Jinu Nyachhyon, Mridul Sharma, Prajwal Thapa +1
The Nepali language has distinct linguistic features, especially its complex script (Devanagari script), morphology, and various dialects,which pose a unique challenge for Natural…