25 citations · 56 across the 16 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.CL2026
In-Place Tokenizer Expansion for Pre-trained LLMs
Jimmy T. H. Smith, Tarek Dakhran, Alberto Cabrera +7
A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflecting the deployment priorities at that time. When those prioriti…
cs.LG2026
MMAI Gym for Science: Training Liquid Foundation Models for Drug Discovery
Maksim Kuznetsov, Zulfat Miftahutdinov, Rim Shayakhmetov +17
General-purpose large language models (LLMs) that rely on in-context learning do not reliably deliver the scientific understanding and performance required for drug discovery tasks…