5 papers
Revisiting Multilingual Data Mixtures in Language Model Pretraining
Negar Foroutan, Paul Teiletche, Ayush Kumar Tarun +1
The impact of different multilingual data mixtures in pretraining large language models (LLMs) has been a topic of ongoing debate, often raising concerns about potential trade-offs…
ModernVBERT: Towards Smaller Visual Document Retrievers
Paul Teiletche, Quentin Macé, Max Conti +4
Retrieving specific information from a large corpus of documents is a prevalent industrial use case of modern AI, notably due to the popularity of Retrieval-Augmented Generation (R…
MMORE: Massive Multimodal Open RAG & Extraction
Alexandre Sallinen, Stefan Krsteski, Paul Teiletche +7
We introduce MMORE, an open-source pipeline for Massive Multimodal Open RetrievalAugmented Generation and Extraction, designed to ingest, transform, and retrieve knowledge from het…
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Project Apertus, Alejandro Hernández-Cano, Alexander Hägele +100
We present Apertus, a fully open suite of large language models (LLMs) designed to address two systemic shortcomings in today's open model ecosystem: data compliance and multilingu…
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ
Marc-Antoine Allard, Matin Ansaripour, Maria Yuffa +1
Large Language Models (LLMs) often struggle with tasks requiring mathematical reasoning, particularly multiple-choice questions (MCQs). To address this issue, we developed LLaMa-Sc…