134 citations · 244 across the 13 of their papers we have counts for
17 papers
L3Cube-MahaSBERT and HindSBERT: Sentence BERT Models and Benchmarking BERT Sentence Representations for Hindi and Marathi
Ananya Joshi, Aditi Kajale, Janhavi Gadre +2
Sentence representation from vanilla BERT models does not work well on sentence similarity tasks. Sentence-BERT models specifically trained on STS or NLI datasets are shown to prov…
Towards Simple and Efficient Task-Adaptive Pre-training for Text Classification
Arnav Ladkat, Aamir Miyajiwala, Samiksha Jagadale +2
Language models are pre-trained using large corpora of generic data like book corpus, common crawl and Wikipedia, which is essential for the model to understand the linguistic char…
A Review of Challenges in Machine Learning based Automated Hate Speech Detection
Abhishek Velankar, Hrushikesh Patil, Raviraj Joshi
The spread of hate speech on social media space is currently a serious issue. The undemanding access to the enormous amount of information being generated on these platforms has le…
L3Cube-MahaNLP: Marathi Natural Language Processing Datasets, Models, and Library
Raviraj Joshi
Despite being the third most popular language in India, the Marathi language lacks useful NLP resources. Moreover, popular NLP libraries do not have support for the Marathi languag…
L3Cube-HingCorpus and HingBERT: A Code Mixed Hindi-English Dataset and BERT Language Models
Ravindra Nayak, Raviraj Joshi
Code-switching occurs when more than one language is mixed in a given sentence or a conversation. This phenomenon is more prominent on social media platforms and its adoption is in…
L3Cube-MahaNER: A Marathi Named Entity Recognition Dataset and BERT models
Parth Patil, Aparna Ranade, Maithili Sabane +2
Named Entity Recognition (NER) is a basic NLP task and finds major applications in conversational and search systems. It helps us identify key entities in a sentence used for the d…