2 papers
cs.CL2026
Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages
Sofiia Riazhskykh, Nam Luu, Ondřej Bojar
Pretraining LLMs on artificial languages ("pre-pretraining") is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33% of training tokens needed t…
cs.CL2025
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
Nam Luu, OndÅej Bojar
Speech Translation (ST) is a machine translation task that involves converting speech signals from one language to the corresponding text in another language; this task has two dif…