2 papers
cs.CL2025
The Lucie-7B LLM and the Lucie Training Dataset: Open resources for multilingual language generation
Olivier Gouvert, Julie Hunter, Jérôme Louradour +6
We present both the Lucie Training Dataset and the Lucie-7B foundation model. The Lucie Training Dataset is a multilingual collection of textual corpora centered around French and…
cs.IR2025
TEARS: Textual Representations for Scrutable Recommendations
Emiliano Penaloza, Olivier Gouvert, Haolun Wu +1
Traditional recommender systems rely on high-dimensional (latent) embeddings for modeling user-item interactions, often resulting in opaque representations that lack interpretabili…