2 papers
cs.CL2026
Understanding Data Temporality Impact on Large Language Models Pre-training
Hippolyte Pilchen, Romain Fabre, Franck Signe Talla +2
Large language models (LLMs) are typically trained on shuffled corpora, yielding models whose knowledge is frozen at train time and whose temporal grounding remains poorly understo…
cs.CL2026
Simultaneous Speech-to-Speech Translation Without Aligned Data
Tom Labiausse, Romain Fabre, Yannick Estève +2
Simultaneous speech translation requires translating source speech into a target language in real-time while handling non-monotonic word dependencies. Traditional approaches rely o…