activity
20242026
collaborators

7 papers

cs.CL2026

HEBATRON: A Hebrew-Specialized Open-Weight Mixture-of-Experts Language Model

Noam Kayzer, Dan Revital, Ori Bar Joseph +10

We present Hebatron, a Hebrew-specialized open-weight large language model built on the NVIDIA Nemotron-3 sparse Mixture-of-Experts architecture. Training employs a three-phase eas…

cs.CL2026

Dicta-LM 3.0: Advancing The Frontier of Hebrew Sovereign LLMs

Shaltiel Shmidman, Avi Shmidman, Amir DN Cohen +1

Open-weight LLMs have been released by frontier labs; however, sovereign Large Language Models (for languages other than English) remain low in supply yet high in demand. Training…

cs.CL2025

Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces

Shaltiel Shmidman, Asher Fredman, Oleg Sudakov +1

Test-time scaling, which leverages additional computation during inference to improve model accuracy, has enabled a new class of Large Language Models (LLMs) that are able to reaso…

cs.CL2025

NeoDictaBERT: Pushing the Frontier of BERT models for Hebrew

Shaltiel Shmidman, Avi Shmidman, Moshe Koppel

Since their initial release, BERT models have demonstrated exceptional performance on a variety of tasks, despite their relatively small size (BERT-base has ~100M parameters). Neve…

cs.CL2025

Splintering Nonconcatenative Languages for Better Tokenization

Bar Gazit, Shaltiel Shmidman, Avi Shmidman +1

Common subword tokenization algorithms like BPE and UnigramLM assume that text can be split into meaningful units by concatenative measures alone. This is not true for languages su…

cs.CL2024

Diversity Over Quantity: A Lesson From Few Shot Relation Classification

Amir DN Cohen, Shauli Ravfogel, Shaltiel Shmidman +1

In few-shot relation classification (FSRC), models must generalize to novel relations with only a few labeled examples. While much of the recent progress in NLP has focused on scal…