14 citations · 22 across the 7 of their papers we have counts for
6 papers · 1 filter
Adapting LLMs to Hebrew: Unveiling DictaLM 2.0 with Enhanced Vocabulary and Instruction Capabilities
Shaltiel Shmidman, Avi Shmidman, Amir DN Cohen +1
Training large language models (LLMs) in low-resource languages such as Hebrew poses unique challenges. In this paper, we introduce DictaLM2.0 and DictaLM2.0-Instruct, two LLMs der…
MRL Parsing Without Tears: The Case of Hebrew
Shaltiel Shmidman, Avi Shmidman, Moshe Koppel +1
Syntactic parsing remains a critical tool for relation extraction and information extraction, especially in resource-scarce languages where LLMs are lacking. Yet in morphologically…
Introducing DictaLM -- A Large Generative Language Model for Modern Hebrew
Shaltiel Shmidman, Avi Shmidman, Amir David Nissan Cohen +1
We present DictaLM, a large-scale language model tailored for Modern Hebrew. Boasting 7B parameters, this model is predominantly trained on Hebrew-centric data. As a commitment to…
DictaBERT: A State-of-the-Art BERT Suite for Modern Hebrew
Shaltiel Shmidman, Avi Shmidman, Moshe Koppel
We present DictaBERT, a new state-of-the-art pre-trained BERT model for modern Hebrew, outperforming existing models on most benchmarks. Additionally, we release three fine-tuned v…
Introducing BEREL: BERT Embeddings for Rabbinic-Encoded Language
Avi Shmidman, Joshua Guedalia, Shaltiel Shmidman +3
We present a new pre-trained language model (PLM) for Rabbinic Hebrew, termed Berel (BERT Embeddings for Rabbinic-Encoded Language). Whilst other PLMs exist for processing Hebrew t…
Shamela: A Large-Scale Historical Arabic Corpus
Yonatan Belinkov, Alexander Magidow, Maxim Romanov +2
Arabic is a widely-spoken language with a rich and long history spanning more than fourteen centuries. Yet existing Arabic corpora largely focus on the modern period or lack suffic…