3 citations · 4 across the 6 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024★ 1 cited
Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages
S. Tamang, D. J. Bora
Large Language Models (LLMs) based on transformer architectures have revolutionized a variety of domains, with tokenization playing a pivotal role in their pre-processing and fine-…
cs.CL2024
Enhancing Assamese NLP Capabilities: Introducing a Centralized Dataset Repository
S. Tamang, D. J. Bora
This paper introduces a centralized, open-source dataset repository designed to advance NLP and NMT for Assamese, a low-resource language. The repository, available at GitHub, supp…
cs.CL2024★ 3 cited
Performance Evaluation of Tokenizers in Large Language Models for the Assamese Language
Sagar Tamang, Dibya Jyoti Bora
Training of a tokenizer plays an important role in the performance of deep learning models. This research aims to understand the performance of tokenizers in five state-of-the-art…