1 paper · 1 filter
Christian Herold, Michael Kozielski, Nicholas Santavas +2
When using an LLM to process text outside the training domain(s), an often overlooked factor is vocabulary mismatch, where the general-domain tokenizer fails to capture frequent do…