7 papers
HEBATRON: A Hebrew-Specialized Open-Weight Mixture-of-Experts Language Model
Noam Kayzer, Dan Revital, Ori Bar Joseph +10
We present Hebatron, a Hebrew-specialized open-weight large language model built on the NVIDIA Nemotron-3 sparse Mixture-of-Experts architecture. Training employs a three-phase eas…
Dicta-LM 3.0: Advancing The Frontier of Hebrew Sovereign LLMs
Shaltiel Shmidman, Avi Shmidman, Amir DN Cohen +1
Open-weight LLMs have been released by frontier labs; however, sovereign Large Language Models (for languages other than English) remain low in supply yet high in demand. Training…
Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces
Shaltiel Shmidman, Asher Fredman, Oleg Sudakov +1
Test-time scaling, which leverages additional computation during inference to improve model accuracy, has enabled a new class of Large Language Models (LLMs) that are able to reaso…
NeoDictaBERT: Pushing the Frontier of BERT models for Hebrew
Shaltiel Shmidman, Avi Shmidman, Moshe Koppel
Since their initial release, BERT models have demonstrated exceptional performance on a variety of tasks, despite their relatively small size (BERT-base has ~100M parameters). Neve…
Splintering Nonconcatenative Languages for Better Tokenization
Bar Gazit, Shaltiel Shmidman, Avi Shmidman +1
Common subword tokenization algorithms like BPE and UnigramLM assume that text can be split into meaningful units by concatenative measures alone. This is not true for languages su…
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
Amir DN Cohen, Shauli Ravfogel, Shaltiel Shmidman +1
In few-shot relation classification (FSRC), models must generalize to novel relations with only a few labeled examples. While much of the recent progress in NLP has focused on scal…